Hoard CTIHoard CTI

Scrapers & Collectors

Most feeds are a scheduled fetch and a field mapping — those are config, not code. Real code is reserved for the awkward sources.

Proposed design

This page describes decisions taken during design. It is not yet implemented — treat it as the intended shape of the system rather than a description of running infrastructure.

Whatever a collector is written in, it does the same job: fetch from a source and emit envelopes onto the queue. It holds no database credentials, knows nothing of the schema, and does not canonicalise anything.

That leaves one real question — how much of a collector needs to be code at all.

Most feeds don't need a scraper

Roughly 80% of CTI sources are the same shape: GET a text, CSV, JSON, or STIX file on a schedule, parse it, map the fields. There is no logic there worth writing three times.

That shape is declarative. A manifest format plus one generic Go collector covers it, and adding a feed becomes a config PR rather than a new service.

Why this matters for an open-source project

A contributor who wants to add a feed shouldn't have to learn the ingest architecture, get database access, or ship a container. They should be able to open a pull request containing a manifest file.

Reserve real code for the awkward sources

The rest of the sources are genuinely awkward — authenticated APIs with unusual pagination, JavaScript-rendered pages, formats that need a specialist parser. Those get written as collectors, and language choice follows capability, not preference.

Go

The manifest executor, the ingest service, and anything on the volume hot path.

Python

Where the ecosystem is: stix2, pymisp, yara-python, and ML-adjacent work.

Node

Playwright for JS-rendered sites, and vendor SDKs that only ship JavaScript.

The cost of three languages

Three languages is a real cost, not a free choice: three dependency trees, three container images, three sets of CVEs to track, and three different ways to write a retry loop.

It is worth paying where the ecosystem genuinely differs — nothing outside Python matches stix2 and pymisp, and nothing outside Node drives a headless browser as well as Playwright.

It is not worth paying to fetch a CSV. That is what the manifest is for.

What every collector owes the pipeline

ResponsibilityDetail
Emit the shared envelopeTypes are generated from the canonical Protobuf / JSON Schema definition
Report raw valuesraw_value as seen at the source — defanged, mixed-case, trailing slashes and all
Identify itselfsource_id and source_run_id, so every record traces back to the run that produced it
Stay out of the databaseNo credentials, no schema knowledge, no writes

Scheduling is still open: independent cron per collector, or a control plane handing out jobs.

On this page