Ingest
One canonical envelope, one process writing to Postgres, and a raw archive written before anything is parsed.
Proposed design
This page describes decisions taken during design. It is not yet implemented — treat it as the intended shape of the system rather than a description of running infrastructure.
Hoard CTI collects from many sources, written in several languages, by many contributors. What holds that together is not a shared codebase — it's centralising the contract and the writer, and nothing else.
Two rules
One canonical envelope
Defined once in Protobuf or JSON Schema, with types generated for all three languages. This schema file is the actual centre of the system.
Exactly one process writes to Postgres
The Go ingest service. Scrapers hold no database credentials and know nothing of the schema.
The second rule is the one that earns its keep. If Go, Python, and Node each implement their own defanging and hash normalisation, the three will drift into subtly different results. The unique constraint on (type, canonical_value) then stops deduplicating — and because nothing errors, it goes unnoticed for months.
The envelope
source_id
source_run_id
collected_at
content_hash
observable { type, raw_value }
context { first_seen, last_seen, confidence, tags, tlp }
raw_ref // pointer to archived payload
schema_versionNote that the field is raw_value, not value.
Scrapers report, they don't interpret
A scraper emitting hxxp://evil[.]com is behaving correctly. Canonicalisation happens exactly once, in the ingest service — see Canonicalisation.
Transport
| Option | When to use it |
|---|---|
| Redis Streams | The default. Redis is already running for IOC lookups, so it adds no new infrastructure, and it provides consumer groups, acks, and replay. |
| NATS JetStream | When several independent consumer groups need the same stream — ingest, enrichment, alerting — or when per-subject retention is needed. |
| Postgres (pgmq / River) | One less system to run, but it puts queue write load on the database the split hot path exists to protect. |
| Kafka / Redpanda | Only at millions of events per minute. |
Archive raw before normalising
Write the untouched payload to R2 keyed by its content hash, put that key in the envelope as raw_ref, and only then parse it.
The cheapest insurance in the design
Parsers improve and feeds change format, and you will want to reprocess months of history without re-scraping. Sometimes re-scraping is impossible — hidden services disappear. This is also the step most often skipped.
Because the archive is keyed by content hash, a re-run that sees identical bytes writes nothing new, and a reprocessing job can replay history through an improved parser without touching the original source at all.
Where enrichment attaches
The enrichment pipeline — parse, research, verify, output — has two plausible attachment points: consuming from the queue as an independent consumer group, or triggering after insert. This is still open, and it is one of the reasons NATS JetStream is worth keeping in view.
Open items
- Envelope schema versioning and migration policy
- Manifest format specification for declarative feeds — see Scrapers & Collectors
- Scraper scheduling: independent cron per collector, or a control plane handing out jobs
- Where the enrichment pipeline attaches