Skip to content
← all work

Netision · 2026 / case 05

Observability, outbound only

Telemetry from machines inside a customer's network, where exactly one process is allowed to talk to the internet.

Designed and built the backends: Go agent, collector, probe, control plane

Go · OpenTelemetry Collector · gRPC + mTLS · BadgerDB · FastAPI · Redpanda · ClickHouse

metrics flowing to ClickHouse, first real pipeline
47k+ in 10 min
commit
processes that egress from the datacenter
1
doc
commits across the five repos
116
count

An on-prem observability platform, built between March and May 2026. A Go agent on each machine, a Go collector per datacenter, a standalone probe, agentless discovery, and a FastAPI control plane on Redpanda and ClickHouse. The console UI was built by teammates.

employer work: written up as architecture and decisions; client names, internal hosts and source stay private.

fig.how it fits together

01/04One static Go binary per machine. Modules (host metrics, logs, SNMP, ICMP) post to an in-process bus, and a single forwarder drains it in batches.

try it

What happens when the uplink drops?

Pick a moment and watch the pipeline handle it: telemetry inside a network that only lets one thing out.

moment · “Every 10 seconds: host metrics from three machines reach the cloud.”

t = 0 ms · timings illustrative
  1. moduleshost metrics · 15 s
  2. forwarder1,000 per batch · gRPC
  3. collectormTLS · written first
  4. uplinkHTTPS :443
  5. storageRedpanda → ClickHouse

starting…

§ 1The motive

Infrastructure teams want monitoring without opening their network. So the design starts from the firewall: one machine needs outbound :443, and every other machine stays locked down.

§ 2Decisions

  • Pure pull. The cloud never connects into the customer network, and the collector never connects to agents. Config arrives because something asked for it.
  • Write-ahead log first, and honest about its gaps. The collector writes every batch to disk and deletes it only when the cloud acknowledges it. What it doesn't do yet: re-send a batch that ran out of retries before the next restart. The agents' own log doesn't replay at all, so a failed push can drop the batch in flight. The code comments say so; the README still promises more, and that's next on my list.
  • Rewrite once the shape was clear. The first prototype wrapped an external collector binary in Python. After two weeks it was deleted for a Go agent that embeds the collector library in-process.
  • Plan, then execute. Each repository started as a plan document, then one commit per step. Working with Claude Code, the control plane went through eight planned phases in a day.

§ 3What this doesn't fix

Keeping control-plane state in ClickHouse is a contested call. Postgres has a real case there (concurrency, revocation lag), and I'd revisit it before the fleet grows.