Scrapers & Collectors
Most feeds are a scheduled fetch and a field mapping — those are config, not code. Real code is reserved for the awkward sources.
Proposed design
This page describes decisions taken during design. It is not yet implemented — treat it as the intended shape of the system rather than a description of running infrastructure.
Whatever a collector is written in, it does the same job: fetch from a source and emit envelopes onto the queue. It holds no database credentials, knows nothing of the schema, and does not canonicalise anything.
That leaves one real question — how much of a collector needs to be code at all.
Most feeds don't need a scraper
Roughly 80% of CTI sources are the same shape: GET a text, CSV, JSON, or STIX file on a schedule, parse it, map the fields. There is no logic there worth writing three times.
That shape is declarative. A manifest format plus one generic Go collector covers it, and adding a feed becomes a config PR rather than a new service.
Why this matters for an open-source project
A contributor who wants to add a feed shouldn't have to learn the ingest architecture, get database access, or ship a container. They should be able to open a pull request containing a manifest file.
Reserve real code for the awkward sources
The rest of the sources are genuinely awkward — authenticated APIs with unusual pagination, JavaScript-rendered pages, formats that need a specialist parser. Those get written as collectors, and language choice follows capability, not preference.
Go
The manifest executor, the ingest service, and anything on the volume hot path.
Python
Where the ecosystem is: stix2, pymisp, yara-python, and ML-adjacent work.
Node
Playwright for JS-rendered sites, and vendor SDKs that only ship JavaScript.
The cost of three languages
Three languages is a real cost, not a free choice: three dependency trees, three container images, three sets of CVEs to track, and three different ways to write a retry loop.
It is worth paying where the ecosystem genuinely differs — nothing outside Python matches stix2 and pymisp, and nothing outside Node drives a headless browser as well as Playwright.
It is not worth paying to fetch a CSV. That is what the manifest is for.
What every collector owes the pipeline
| Responsibility | Detail |
|---|---|
| Emit the shared envelope | Types are generated from the canonical Protobuf / JSON Schema definition |
| Report raw values | raw_value as seen at the source — defanged, mixed-case, trailing slashes and all |
| Identify itself | source_id and source_run_id, so every record traces back to the run that produced it |
| Stay out of the database | No credentials, no schema knowledge, no writes |
Scheduling is still open: independent cron per collector, or a control plane handing out jobs.