Your Logs Were Never Ours to Keep


When you deploy something to Brimble, we collect your logs so the dashboard can show them back to you.
That is useful. It is also not enough.
At some point, logs stop being a thing you want to look at in a hosting dashboard. They become part of your incident response, your compliance story, your search workflow, your alerting setup, or your weird homegrown ClickHouse thing that nobody is allowed to delete because it answers one query the business cares about.
So the customer problem was simple:
“I can see my logs in Brimble. Can I send them somewhere else?”
Now you can.
Brimble Log Drains forward a project’s logs in real time to any HTTPS endpoint or any S3-compatible bucket. Datadog, Axiom, Better Stack, R2, MinIO, AWS S3, your own ingestion service behind a very judgmental nginx config — all fine.
We send gzipped NDJSON. There are three sources:
- app logs from container stdout/stderr
- system events like deploys, crashes, OOM kills, and scaling
- HTTP request logs from the edge
It is included on paid plans. No per-GB fee. Vercel charges $0.50/GB for log drains. Railway does not offer drains directly; you run a sidecar. We wanted the boring version: turn it on, choose a destination, receive logs.
The implementation is less glamorous than that. Which is good.
We did not build a log pipeline.
The obvious thing we did not do
Brimble already had logs.
The old path looked roughly like this:
container stdout/stderr
↓
Promtail
↓
Loki
↓
Brimble dashboardSo the first idea was obvious: read logs back out of Loki and forward them to customer destinations.
That idea lasted maybe ten minutes.
Nothing really wants to read logs out of Loki. The ecosystem is full of tools that push logs into it. Promtail, Vector, Alloy, Fluent Bit, random agents with YAML crimes. They all go in the same direction.
Reading out of Loki would have meant building a polling system. Now we have cursors. Now we have backoff. Now we have per-customer delivery state. Now we have to handle partial failures. Now we have to explain why the customer got a duplicate log line during a deploy because a cursor update raced a retry.
This is the part where a simple feature becomes a distributed systems hobby.
We rejected it.
The second idea was to build a delivery service. A Go daemon that reads from Docker, tails edge logs, batches events, retries delivery, writes buffers to disk, exposes Prometheus metrics, reloads config, handles S3 uploads, handles HTTP drains, and slowly becomes a worse version of a project maintained by people who already spent years thinking about this.
We rejected that too.
Vector already exists.
Vector is Datadog’s open-source Rust log shipper. It can read Docker logs. It can tail files. It can accept HTTP events. It has per-sink disk buffers, batching, retries, adaptive concurrency, compression, S3 support, HTTP support, and Prometheus metrics.
That is the whole delivery engine.
Our “log drain engine” is a generated Vector config file.
The shape of the system
Brimble runs on bare metal with Nomad, Consul, and Vault.
Nomad schedules workloads. Consul stores dynamic config. Vault stores secrets. We already use this pattern in other parts of the platform, so log drains became another generated config problem.
A customer creates a drain.
We store the non-secret drain config in Consul KV.
The secret bits go into Vault.
Then Nomad’s template runner renders vector.toml.
Vector runs with:
vector --config "/local/vector/vector.toml" --watch-configSo when a drain changes, Vector does not need a deploy. It does not need a restart. It does not even need to know Consul exists. It just watches a file.
That last part matters. Systems get weird when every process talks to every control plane. Vector should ship logs. It should not become a Consul client, a Vault client, and a Brimble API client. Nomad already knows how to render templates from Consul and Vault. Let the thing that is good at the thing do the thing.
The render loop looks like this:
The important trick: Vector never talks to Consul or Vault. Nomad renders a file, and Vector watches it.
A customer creating a drain becomes a live Vector pipeline in seconds, with zero deploys and zero restarts.
A trimmed version of the config looks like this:
data_dir = "/alloc/data"
[api]
enabled = true
address = "0.0.0.0:8686"
[sources.system_events]
type = "http_server"
address = "0.0.0.0:8687"
decoding.codec = "json"
[sources.docker]
type = "docker_logs"
include_labels = [
"project",
"identifier",
"job_name",
"com.hashicorp.nomad.alloc_id",
]
[transforms.envelope_app]
type = "remap"
inputs = ["docker"]
source = '''
.project = .label."project"
.slug = .label."identifier"
.service = .label."job_name"
.instance = .label."com.hashicorp.nomad.alloc_id" || .container_name
.region = "eu-central"
.source = "app"
.event = null
.http = null
if !exists(.timestamp) {
.timestamp = now()
}
.level = null
parsed, err = parse_json(.message)
if err == null && is_object(parsed) && exists(parsed.level) {
.level = to_string(parsed.level) ?? null
}
del(.label)
del(.container_created_at)
del(.container_id)
del(.stream)
del(.host)
del(.image)
'''
[transforms.envelope]
type = "filter"
inputs = ["system_events", "envelope_app"]
condition = 'exists(.project) && !is_nullish(.project)'Every event gets normalized into the same envelope. App logs, system events, and edge request logs all become one shape.
Then each drain renders as a filter plus a sink.
For an HTTP drain:
[transforms.route_drn_123]
type = "filter"
inputs = ["envelope"]
condition = '.project == "prj_abc" && includes(["app", "system", "http"], .source)'
[sinks.drain_drn_123]
type = "http"
inputs = ["route_drn_123"]
uri = "https://logs.example.com/ingest"
encoding.codec = "json"
framing.method = "newline_delimited"
compression = "gzip"
batch.timeout_secs = 1
batch.max_bytes = 1048576
request.timeout_secs = 10
request.retry_attempts = 5
buffer.type = "disk"
buffer.max_size = 268435456
buffer.when_full = "drop_newest"
[sinks.drain_drn_123.request.headers]
"User-Agent" = "Brimble-Log-Drain/1.0"
"Brimble-Drain-Id" = "drn_123"
"Content-Type" = "application/x-ndjson"
"Content-Encoding" = "gzip"
"Authorization" = "Bearer ..."For S3:
[sinks.drain_drn_456]
type = "aws_s3"
inputs = ["route_drn_456"]
bucket = "my-log-bucket"
region = "auto"
endpoint = "https://example.r2.cloudflarestorage.com"
key_prefix = "brimble/prj_abc/dt=%Y-%m-%d/"
filename_time_format = "%s-%N"
filename_append_uuid = true
filename_extension = "ndjson.gz"
encoding.codec = "json"
framing.method = "newline_delimited"
compression = "gzip"
batch.timeout_secs = 300
batch.max_bytes = 10485760
buffer.type = "disk"
buffer.max_size = 268435456
buffer.when_full = "drop_newest"
[sinks.drain_drn_456.auth]
access_key_id = "..."
secret_access_key = "..."That is basically it.
There is more template glue around node types, Vault reads, and Consul iteration, but the core idea is boring: produce one event stream, filter it per drain, send it to a sink.
The boringness is load-bearing.
One system job, different nodes
We run Vector as a Nomad system job.
That means one copy per eligible node: runners, database nodes, and load balancers.
The rendered config differs by node type.
Runner and database nodes use Docker logs:
[sources.docker]
type = "docker_logs"Load balancers tail Caddy access logs:
[sources.caddy_access]
type = "file"
include = ["/var/log/caddy/*.log"]
read_from = "end"
ignore_older_secs = 86400Caddy writes one access log file per project:
/var/log/caddy/<project-id>.logSo the edge path does not trust a user-supplied project field. The project comes from the filename.
The transform turns Caddy’s JSON into the same envelope:
path = string!(.file)
line = parse_json!(string!(.message))
if line.request.uri == "/__brimble/health" {
abort
}
m = parse_regex!(path, r'/var/log/caddy/(?P<project>[^/]+)\.log$')
ev = {}
ev.timestamp = from_unix_timestamp!(
to_int(float!(line.ts) * 1000.0),
unit: "milliseconds"
)
ev.project = m.project
ev.service = "web"
ev.source = "http"
ev.instance = "edge"
ev.region = "eu-central"
ev.http = {
"method": line.request.method,
"path": line.request.uri,
"status": line.status,
"duration_ms": round(float!(line.duration) * 1000, precision: 2),
"bytes": line.size,
"host": downcase(string!(line.request.host))
}
ev.message = join!([
string!(line.request.method),
string!(line.request.uri),
to_string!(line.status)
], " ")
. = evI like this part because the isolation boundary is physical and dumb. The file is named after the project. Docker containers get labels from the scheduler. Users do not get to invent their own tenancy metadata and hope we believe them.
Hope is not an isolation primitive.
Per-drain isolation
A log drain is best-effort delivery.
That is intentional.
If your endpoint is down, Brimble should not slow down your app. If your S3 bucket starts returning 403s, it should not affect another customer’s drain. If Datadog has a bad day, my control plane should not be pulled into the group project.
Each rendered drain has its own Vector sink and its own disk buffer:
buffer.type = "disk"
buffer.max_size = 268435456
buffer.when_full = "drop_newest"When the destination fails, that drain’s buffer fills. Vector retries. If it stays broken, new events for that drain are dropped. Other drains keep moving.
This is the important tradeoff: log drains are not a durable queue.
They are a delivery pipe with bounded buffers. We could build stronger durability later, but then we are back to owning a delivery engine, storage semantics, replay controls, customer-visible cursor state, and a lot of meetings that should have been SQL migrations.
For v1, bounded best-effort delivery is the right contract.
We also run a small Go reconciler that polls Vector’s per-sink Prometheus metrics. It drives a simple state machine:
ok → failing → auto-disabledWhen a drain fails continuously, we email the customer. After 24 hours of continuous failure, we auto-disable it.
A broken destination should be visible. It should not become a permanent disk pressure machine.
SSRF: “send logs to my internal service” is not a feature
HTTP drains accept a URL.
That is convenient and dangerous.
Our nodes sit on private networks. Some are on a Tailscale mesh. Without protection, a customer could create a drain pointed at an internal address and use Vector as a very strange log-powered SSRF tool.
That sounds fake until you build anything that accepts webhooks.
So drain URLs go through SSRF checks before we write config.
We resolve the hostname and reject private ranges. That includes the usual suspects, and also 100.64.0.0/10, because Tailscale uses carrier-grade NAT space and “drain your logs to my Vault instance” is not on the roadmap.
The app-level check is not the only line of defense. We also use nftables egress rules on the nodes.
The rule of thumb is simple: validation is nice, packet filters are nicer.
The part where Vector taught us humility
A lot broke while building this.
The funniest class of bugs was “this config validates, then dies in production.”
Vector’s VRL is good, but it is very particular. At one point every Vector task crash-looped with exit 78, EX_CONFIG. The process booted, read the config, and died in about a second.
The cause was this tiny thing:
.instance = .label."com.hashicorp.nomad.alloc_id" ?? .container_nameThat looks reasonable if you write JavaScript or TypeScript all day.
In VRL, ?? is error-coalescing. The left side was an infallible field access in that context, so Vector rejected it with E651. What I wanted was null coalescing:
.instance = .label."com.hashicorp.nomad.alloc_id" || .container_nameThen there was this:
.timestamp = to_timestamp!(line.ts)Also wrong. In Vector 0.49, that function is gone. The fix was:
.timestamp = from_unix_timestamp!(
to_int(float!(line.ts) * 1000.0),
unit: "milliseconds"
)Then we had silent event loss. This one was rude.
We originally wrote:
.level = string!(parsed.level)That asserts the value is already a string. Plenty of loggers emit numeric levels. Pino and Bunyan use values like 30.
So a perfectly valid log line with "level": 30 could abort the transform and drop the whole event.
The fix:
.level = to_string(parsed.level) ?? nullCoerce. Do not assert. Assertions are where log lines go to die.
The other trap was validation. vector validate --no-environment passed configs that crashed at runtime, because without the right environment it did not build the same sources, so some VRL was never type-checked.
The lesson: validate like production. Give Vector a real data directory. Run it where the Docker socket exists. Use --deny-warnings.
Static config validation is not magic. It is just a test with better branding.
The Vault bug that looked like a missing secret
The deepest rabbit hole was Vault.
Vector config is rendered by Nomad templates. Drain secrets live in Vault. Easy enough.
Then templates started failing with Missing.
The secret existed.
The path was right.
The policy looked right.
Still Missing.
This is one of those bugs where the error message is technically correct in a way that makes you worse at your job. In a Nomad template, a denied Vault read can appear as Missing. It does not necessarily mean 404. It can mean 403.
This was a permissions mismatch, not a missing secret. The token rendering the config was not being granted the policy we thought it was. Once the policy binding was corrected, the reads went through.
The lesson is the useful part, and it is not Vault-specific. When a secret is "missing" but you can see it sitting right there, stop trusting the error string and go read the audit trail. A denied read hides in plain sight. Computers remain undefeated.
Product shape matters too
There was also a product modeling mistake.
The first implementation treated drains as project-local config. That made sense for the first screen and became awkward almost immediately.
Real teams do not want to paste the same Datadog endpoint into seven projects. They want one account/team-level drain config and then a per-project toggle.
So we refactored:
team drain config
↓
project subscription
↓
Consul-rendered Vector routeDeleting a project removes that project’s subscriptions and Consul entries. It does not delete the global drain config.
Downgrading a plan disables project subscriptions, but keeps the global configs intact.
Deleting a drain is a hard delete. We remove the drain record, project subscriptions, Consul entries, and Vault secrets. For Brimble-managed S3 buckets, we also revoke the scoped storage credential when the drain is deleted or replaced.
Soft delete was the wrong behavior here. A deleted log destination should not haunt the control plane.
What we punted on
Ordering across nodes is not guaranteed.
If your app runs on three machines, those machines ship logs independently. Within a sink, Vector batches. Across the fleet, clocks and networks do what clocks and networks do. Every event has a timestamp. Use that.
Backfill is not included.
A drain starts forwarding after it is created. We do not replay historical Loki data into your new drain. That would require the polling/cursor system we very deliberately avoided.
Syslog is not included.
HTTPS and S3-compatible buckets cover the destinations we care about for v1. Syslog can come later if enough customers ask for it and I lose an argument with myself.
Exactly-once delivery is also not included, because this is a log drain, not a bank ledger.
Where to find it
Log Drains are available on paid Brimble plans now.
Open Settings and go to the Log Drains tab, then click Add drain. Pick an HTTPS endpoint or an S3-compatible bucket, and choose the sources you want: app logs, system events, HTTP request logs. A drain is created once, at the account level. Then, from any project’s Log Drains page, enable that drain to start forwarding the project’s logs to it — a project only sends logs to the drains it is subscribed to.
A drain is configured once at the team level, then projects subscribe to it.
We send gzipped NDJSON. We do not charge per GB. Broken drains are isolated, monitored, and eventually auto-disabled after 24 hours of continuous failure.
Your logs were never ours to keep.
Now they have somewhere else to go.
Written by
