Conversation
Pivots an Organization to the news coverage that mentions it, emitting one Website per article and relating it back with MENTIONED_IN. The existing Organization enrichers are registry and infrastructure pivots (Whoxy, SIRENE, asnmap), so none of them surface what has been said about an organization publicly. Coverage names the people, partners and incidents that no registry records. Language and market are pinned by default, since the vertical otherwise infers them from where the call originates and answers an English query with articles in another language. Recency is exposed as a time_range param.
The frontend needs a flag for whether an enricher takes params at all. Follows the existing override in ip/to_asn.py, email/to_leaks.py and six other enrichers, matching the fix already applied to reconurge#223.
Author
|
Pulling your Added in |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds
org_to_news, an Organization to Website enricher that asks a searchindex for news coverage of an organization. It runs the name as an
exact-phrase query against the news vertical of the Serply API and emits one
Website node per article, related back to the input Organization with
MENTIONED_IN.Why
Third and last of the three enrichers from #220, after #223 (domain) and #225
(individual).
The three Organization enrichers today are
org_to_domains(Whoxy),org_to_infos(SIRENE) andorg_to_asn(asnmap). All three are registry orinfrastructure pivots: they tell you what the organization filed and what it
runs. None of them tell you what has been said about it. Coverage is often
where the next pivot actually comes from, because an article names the people,
the partners, the acquisition, the breach and the locations that no registry
records, and it is the one source that exists for an organization too small or
too foreign to appear in SIRENE.
@dextmorgn this completes the set you green-lit on #220. All three are
independent PRs and can be merged, reordered or dropped one at a time.
Notes
SERPLY_API_KEYis declared as arequired
vaultSecret, so the enricher is inert until a key exists in theVault, same as
domain_to_dehashedandip_to_fraudscore.endpoint as feat(enrichers): add domain_to_indexed_pages enricher backed by Serply #223 and feat(enrichers): add individual_to_mentions enricher backed by Serply #225, with the vertical selected by Google's own
tbmparameter.
MENTIONED_INis the same label feat(enrichers): add individual_to_mentions enricher backed by Serply #225 introduces, for the same reason: anarticle covering an organization mentions it, it is not a site the
organization owns, so this is not the
HAS_WEBSITEedgedomain_to_websitewrites. That is the one cross-PR coupling here. If you would rather the label
were named something else, say so on either PR and I will change both. If
only one of the two merges the label still works, since neither reads the
other's edges.
en/us. This is notboilerplate: left unset, the vertical infers them from where the call
originates, and a server is not a place. An English-language query for a
well-known company came back as ten Japanese-language articles plus one row
whose
linkwas a thumbnail image rather than an article. Withhlandglset, all ten rows were articles. Both are exposed as params so a French
investigation can ask for
fr/fr.time_rangemaps onto the index's own recency filter (tbs=qdr:d|w|m|y),which matters more here than on a web search: "the last week of coverage" is
usually the question. An unrecognised value warns and searches all of time
rather than failing the scan.
one, but
Websitehas no date field, and I would rather not stuff it intodescriptionor invent a field inflowsint-typesinside a PR that is aboutan enricher. Same for
Website.domain, which I leave towebsite_to_domain. Both are a small follow-up if you want them, and addinga
published_attoWebsiteis the kind of change I would rather youdecided than discovered.
domain/to_dns.pyandip/to_fraudscore.py: registrationvia the decorator, no
registry.pyedit, per-item try/except,Logger.erroron every failure path. The source name rides to
postprocesson a temporaryattribute and is deleted there, the way
domain_to_dnsthreads its sourcedomain, because one scan can cover several organizations and the results
cannot be zipped back to their input.
Organization.nameis typedAny, so it can arrive empty or unset from anupstream pivot. It is coerced and checked before anything is sent, and an
organization with no usable name is skipped with a warning, so it costs no
credit. Pinned by a test over
""," "andNone.ip_to_portscarriessix, so it is not unprecedented. Four of the five are optional with
defaults.
startis the only offset the APIhonours,
numis an approximate cap of about ten rows per call, and a windowthat adds no new link ends the loop instead of spending a credit per
duplicate window. Surplus rows are trimmed against
max_results(default20).
flowsint-enrichers/tests/enrichers/test_org_to_news.py, HTTP layermocked, no key and no network needed. Docs row added in
docs/sources/available-enrichers.mdx, under### Organization, so it doesnot collide with the
### Individualrow feat(enrichers): add individual_to_mentions enricher backed by Serply #225 adds to the same file.Testing
55 is the whole
flowsint-enricherssuite, 32 before this branch plus the 23new ones. I did not run the frontend half of
make lint, since no frontendfile changed here.
Live run, real API key against a local Neo4j 5,
max_results12 andtime_range"m" on "Mistral AI". 12 is over one window, so the second page wasfetched and merged:
Every row was inside the one-month window the run asked for, and title and
snippet land on the node:
Size
3 files, +598. This is the largest of the three and I would rather explain the
number than hide it. 291 lines are the test file and 3 are the docs row, so the
enricher itself is 304. It is longer than #225's 260 because the news vertical
has three knobs a plain web search does not (vertical selection, recency,
locale), and each one carries its schema entry, its mapping and its test.
For reference: #225 was +514, #223 +416, and the merged
domain_to_dns(#182)+336 and #183 +305 across the same three kinds of file.
Happy to cut the test file down if you would rather review it lighter. The
tests I would defend keeping are the locale one and the
time_rangeones,since both pin behavior that is invisible in the output until it is wrong, and
the one that pins the source-name threading across several organizations in one
scan.
Disclosure
I work with Serply. Happy to adjust scope, naming, or drop this entirely if it
is not a direction you want for the project.