All posts
Engineering11 min read

We Stopped Pretending S3 Was a Disk

Ilerioluwa David
Ilerioluwa David@pipe_dev
We Stopped Pretending S3 Was a Disk

Containers are temporary.

That is one of their best features. Nomad can destroy a container, recreate it on another server, and keep the application running.

Files are less enthusiastic about this arrangement.

When a Brimble workload moves from one runner to another, the replacement container needs to find the same persistent data waiting for it. This matters for applications like WordPress, where uploads, themes, plugins, and runtime-generated files all live inside /var/www/html.

For a long time, our answer was S3.

More accurately, our answer evolved through several increasingly convincing ways of making S3 behave like a filesystem:

GeeseFS
   ↓
JuiceFS + S3
   ↓
CephFS

None of them were mistakes. They solved different versions of the same problem.


First, we mounted the bucket

Our original persistent storage layer used GeeseFS:

Container → mounted directory → S3 bucket

Applications wrote to a normal path. GeeseFS translated those filesystem operations into object-storage requests.

This gave us inexpensive, durable storage without operating storage servers. More importantly, it allowed us to ship.

But object storage and filesystems do not think alike.

A filesystem has directories, permissions, renames, metadata, locks, and thousands of tiny mutations.

Object storage has objects and prefixes.

A mount layer can translate between the two, but the translation becomes visible under metadata-heavy workloads. A directory scan may become a series of object listings. A harmless-looking du -sh can suddenly involve much more work than expected.

GeeseFS was useful, but we were asking a bucket to perform an increasingly elaborate disk impression.


JuiceFS gave the bucket a brain

Our next iteration was JuiceFS.

JuiceFS kept object storage as the durable data layer, but introduced a separate metadata system:

This made directory and metadata operations much more predictable. JuiceFS also had a CSI driver, which integrated naturally with our Nomad cluster.

For quite some time, this was our production answer.

We did not move because JuiceFS was broken.

We moved because persistent volumes became more important to Brimble.

The original question was:

How can we cheaply attach durable storage to a container?

Eventually, the question became:

What should sit at the centre of Brimble’s storage platform?

Those are not the same question.


Why we decided to own the storage path

With JuiceFS, one filesystem operation could involve the workload, the CSI client, the metadata database, the object-storage provider, and the network paths connecting them.

That is a reasonable architecture, especially when reducing operational work is the priority.

But we wanted a more direct failure model.

We wanted to see the replicas. We wanted known storage capacity. We wanted native filesystem semantics backed by disks. We wanted customer quotas that mapped cleanly to independent volumes.

Mostly, we wanted persistent volumes to feel less like an integration between several systems and more like storage.

So we bought three servers and installed Ceph.


Three modest servers

Each storage server has roughly:

6 vCPU
18 GB RAM
1 TB SSD
600 Mbit/s network

Our existing machines continue running customer workloads, while the new servers handle storage.

There was one immediate problem: Ubuntu occupied almost the entire 1 TB disk on every server.

We had successfully purchased storage servers without empty storage devices.

Fortunately, the operating system was using only a few gigabytes. We booted each server into rescue mode, shrank the root partition, and reserved the remaining space for Ceph.

The final layout looked roughly like this:

sda      1000G
├─sda1   150G   ext4   /
├─sda2   849G          ceph-osd
└─boot partitions

Because the Ceph partition lived on the same virtual disk as the operating system, we exposed it through LVM:

pvcreate /dev/sda2
vgcreate ceph-osd-vg /dev/sda2
lvcreate -l 100%FREE -n osd0 ceph-osd-vg

Each host now contributed one OSD.

It is not the kind of storage fleet that gets its own keynote, but it has the correct shape.


Ceph over Tailscale

The storage servers do not share a private datacentre network, so they communicate over Tailscale.

We confirmed that the connections were direct rather than relayed through DERP.

The network is not spectacular. Throughput between some nodes ranges from roughly 100 to 250 Mbit/s, and the links are not perfectly symmetrical.

That means normal application storage works, but large recovery operations will be slow. A full rebuild may take hours.

This is one of those startup infrastructure decisions where the architecture is correct, even if the hardware is not yet luxurious.

For us, the first obvious future upgrade is the network.


CephFS in one minute

CephFS splits filesystem work into two main paths.

When an application opens a file, the Metadata Server (MDS) handles directories, permissions, filenames, and file locations. The application then reads or writes the actual file contents directly against the Object Storage Daemons (OSDs).

Application
    │
    ├── metadata requests ──► MDS
    │
    └── file data ──────────► OSDs

The Monitors (MONs) are the cluster’s source of truth: they keep quorum, hand out the cluster map, and decide who is allowed to join. Without enough monitors agreeing with each other, Ceph will not trust the cluster enough to serve traffic safely.

The Managers (MGRs) sit beside the monitors and handle the operational side of the house — status endpoints, dashboards, and the Prometheus metrics we scrape for health and capacity.

The important part is that file data does not flow through the MDS. The MDS manages the namespace, while the OSDs store and replicate the actual data.


Bootstrapping the cluster

We used Ceph Tentacle and cephadm.

Tailscale interfaces use /32 addresses, which cephadm’s normal subnet validation did not enjoy. We skipped monitor-network discovery and supplied the first monitor address directly:

cephadm bootstrap \
  --mon-ip <storage-1-tailnet-ip> \
  --skip-mon-network \
  --skip-monitoring-stack

We then added the other monitors using their own Tailscale addresses.

The finished cluster had:

3 MONs
2 MGRs
2 MDS daemons
3 OSDs

Three monitors allow the cluster to lose one MON and retain quorum. One MDS is active while another waits to take over.

The OSD layout is deliberately simple:

root default
├── host storage-1
│   └── osd.0
├── host storage-2
│   └── osd.1
└── host storage-3
    └── osd.2

Every replica can live on a separate storage host.


Creating CephFS

We created a filesystem named brimblefs.

Its data and metadata pools use three replicas, with a minimum of two:

size     = 3
min_size = 2

Every object exists on all three storage servers. One server can disappear and the filesystem remains available.

The trade-off is capacity.

Roughly 2.5 TiB of raw storage becomes about 806 GiB of usable CephFS capacity.

Replication remains extremely committed to arithmetic.


A green cluster is not a product

Eventually, Ceph reported:

HEALTH_OK

The final state was:

3 MONs in quorum
3 OSDs up and in
1 active MGR and 1 standby
1 active MDS and 1 standby
all placement groups active+clean

That was encouraging, but the real question was whether an application could move between Brimble runners and retain its files.

So we created a quota-controlled CephFS subvolume and gave it a restricted Ceph identity.

The model is simple:

We mounted the same volume on two runners.

A file written on one appeared on the other.

The filesystem worked.

Then we tested something less polite.


The WordPress test

WordPress is an excellent persistent-volume test because it behaves like a real application rather than a benchmark.

It creates files, installs plugins, changes themes, stores uploads, and expects ordinary Linux filesystem semantics.

We mounted the CephFS volume at /var/www/html and started WordPress on one runner.

After completing the installation and uploading content, we deleted the container entirely.

We then started a fresh WordPress container on another runner, attached the same CephFS-backed directory, and pointed it at the same database.

The site returned with the same uploads, themes, and plugins.

WordPress on runner-1
          │
          ▼
container removed
          │
          ▼
WordPress starts on runner-2
          │
          ▼
same filesystem

That was the actual success condition.

Nomad could move a workload without its persistent data belonging to the machine it left behind.


Consul knows where the monitors are

We did not want every CSI job to carry a static list of Ceph monitor addresses.

The Ceph servers already run Consul, so each monitor registers as the service ceph with the tag mon.

Nomad renders the Ceph-CSI configuration from Consul service discovery:

{{ range service "mon.ceph" }}
"{{ .Address }}:{{ .Port }}"
{{ end }}

If the monitor list changes, Nomad rerenders the plugin configuration.

The Ceph cluster ID remains stable and is stored in Consul KV.

This gives us a straightforward regional model:

EU region    → EU Ceph cluster
US region    → US Ceph cluster
Asia region  → Asia Ceph cluster

The cluster ID is not CSI topology. It is simply the stable identity of the storage cluster serving that region.


Turning it into a Nomad primitive

Manual mounts proved the storage path.

Ceph-CSI turns it into a platform feature.

The integration consists of two Nomad jobs:

plugin-cephfs-controller
plugin-cephfs-node

The controller handles volume lifecycle operations. The node plugin runs on each eligible Nomad client and performs the actual mounts.

A simplified Brimble volume request looks like:

id           = "customer-volume-id"
type         = "csi"
plugin_id    = "cephfs0"
capacity_min = "10GiB"
capacity_max = "10GiB"

capability {
  attachment_mode = "file-system"
  access_mode     = "multi-node-multi-writer"
}

parameters {
  clusterID = "regional-ceph-cluster-id"
  fsName    = "brimblefs"
  mounter   = "kernel"
}

Ceph-CSI creates the subvolume and applies the requested capacity as a filesystem quota.

Nomad can then mount that volume on any eligible runner.

No local-disk attachment.

No bucket pretending to be POSIX.

Just a filesystem.


The small runbook

Most first-line checks fit into four commands:

ceph -s
ceph osd tree
ceph fs status
nomad plugin status cephfs0

The output we want is boring:

HEALTH_OK
3 OSDs up and in
all PGs active+clean
1 active MDS and 1 standby

Boring is the desired storage feature.


Telemetry, or: knowing before customers do

A green ceph -s is comforting when you are already looking at it.

It is less useful at 3 a.m., when nobody is looking at it.

So the cluster reports into the same Prometheus setup we use for the rest of Brimble.

Ceph’s manager exposes metrics over its Prometheus module. Our scrapers pull them on a regular interval, and the usual operational picture shows up alongside the rest of the platform:

cluster health
OSD up / in
placement group state
usable capacity
used capacity
MDS active / standby

Health answers the obvious question: is the filesystem still a filesystem, or has it become a science project?

Capacity answers the quieter one: how much room is left before the next customer volume stops being creatable, or before recovery has nowhere to put a replica?

We alert on the boring failures.

A monitor leaving quorum.

An OSD dropping out.

PGs stuck in a state that is not active+clean.

Free space shrinking past the point where growth or recovery still feels comfortable.

The point is not a dashboard full of graphs for their own sake.

The point is that storage stops being a box we only open when something is already broken.

If the cluster is going to hold customer files, it has to argue for attention the same way every other production dependency does: with metrics, not hope.


What Ceph does not solve

Replication is not a backup.

If a disk dies, we have replicas.

If a storage server disappears, the cluster keeps running.

If somebody deletes the wrong customer volume, Ceph will faithfully replicate that deletion across every server.

We still need snapshots, off-cluster backups, retention policies, and restore drills.

Replication protects against infrastructure failure.

Backups protect against people.

We remain the more unpredictable component.


The storage stack grew up with us

Our persistent storage journey now looks like:

GeeseFS
   ↓
JuiceFS + S3
   ↓
CephFS

None of the previous choices were mistakes.

GeeseFS helped us ship quickly.

JuiceFS gave us stronger filesystem semantics without requiring us to operate storage servers.

CephFS arrived when persistent volumes became important enough that we wanted to control the complete storage path.

Infrastructure technologies rarely defeat one another in a dramatic final battle.

Usually, the product changes. The constraints change. The question changes.

For Brimble, the question eventually became:

Can a customer’s application wake up on a different machine and find its filesystem exactly where it left it?

The answer is now yes.

The container can disappear.

The runner can change.

The files do not care.

Written by

Ilerioluwa David
Ilerioluwa David@pipe_dev