Protect ClickHouse
Backups to your own bucket, Marks, weekly Proof, Rewind copies and Pulse for ClickHouse 25.8, 26.3 and 26.8 on your own server or in Docker. Restores go to a backup or a Mark.
Rowsafe protects ClickHouse 25.8, 26.3 and 26.8 (the LTS releases) much like it protects PostgreSQL: backups go to your bucket, encrypted on your server with a passphrase only you hold; a weekly Proof restores a copy and checks it; Rewind restores a copy next to production and brings deleted rows back; Pulse watches the server.
One difference: no restore to any second
ClickHouse keeps no log of changes that Rowsafe could copy, so you can't go back to any second. You go back to a backup or a Mark. Rowsafe takes a full backup every week and a backup of what changed every hour, and a Mark takes one on the spot. So without a Mark, you lose at most about an hour.
Before you start
- ClickHouse runs on Debian 12/13 or Ubuntu 22.04/24.04 (from ClickHouse's own packages, as
clickhouse-server.service), or in the officialclickhouse/clickhouse-serverDocker image. - Its HTTP interface is on (port
8123by default). Rowsafe talks to ClickHouse there, on127.0.0.1. - The
clickhouseprogram is on the server. It comes with ClickHouse's server package. Proof and Rewind copies run it; backups work without it. - You have a bucket and an encryption passphrase, as for PostgreSQL. See Adopt an existing database for the one-command install.
Turn on backups
Run the install command on the server and approve the server in your browser when it prints the link:
curl -fsSL https://rowsafe.sh | sudo shThe installer finds ClickHouse next to (or instead of) PostgreSQL. Nothing restarts:
Rowsafe's own ClickHouse user. Rowsafe needs a user, rowsafe, to take backups and read the server's health. The installer adds it as /etc/clickhouse-server/users.d/rowsafe.xml (owned by root, readable by ClickHouse only). The file holds only a hash of a random password, and the user can only connect from the server itself. ClickHouse loads it by itself within seconds, without a restart. The password is saved for the agent only (/var/lib/rowsafe/engines/clickhouse/logins, readable by the agent alone).
If that isn't possible (ClickHouse's users come from elsewhere), an administrator signs in once instead, and Rowsafe creates the same user with SQL. That password is used once and never saved.
The plan. Like for PostgreSQL, you see what Rowsafe will do and say yes. Nothing in ClickHouse's settings changes. Rowsafe checks that a backup reaches your bucket and opens with your key, and takes the first full backup.
On a server without PostgreSQL, the agent runs as its own system user, rowsafe.
Without a terminal, --protect NAME works for ClickHouse too. When the users file can't be used, it takes an administrator from ROWSAFE_CLICKHOUSE_ADMIN_USER and ROWSAFE_CLICKHOUSE_ADMIN_PASSWORD (used once, never saved).
What Rowsafe does
| How | |
|---|---|
| Full backup | ClickHouse's own BACKUP of every database except the system ones. Every week by default. |
| Differential backup | A BACKUP that copies only the data parts written since the last full backup. ClickHouse never changes a part once it is written, so this is small and quick. Every hour by default. |
| Encryption | ClickHouse writes its backup to a small S3-compatible service inside the agent, which lives only for the length of the backup. The agent encrypts every file, and its name, before it goes to your bucket: the bucket shows when each backup was taken, never your databases' or tables' names. ClickHouse never sees your bucket's keys or your passphrase. |
| Marks | A differential backup taken on the spot and named after the Mark. Restoring to a Mark brings back exactly that moment. |
| Restore | To a backup or a Mark. Choosing a time picks the newest backup that finished at or before it. |
| Proof | Weekly: restore the newest backup into a temporary ClickHouse on the same server (listening on 127.0.0.1 only, with its own data folder), check that every database and table came back with the rows it had, then delete it. |
| Rewind | Restore a copy as it was at a backup or a Mark, next to production, on the same kind of temporary ClickHouse. Compare a table with production and bring missing rows back. ClickHouse tables have no unique key, so rows are compared on all their columns: a row is missing when production has fewer copies of it than the copy. |
| Pulse | Queries, inserts, merges, memory, disk, data parts per partition, mutations, replication, broken parts and the queries running now. Stop query ends one (KILL QUERY) after checking it is still the same query. Cancel mutation stops an ALTER ... UPDATE/DELETE that can't finish (KILL MUTATION); parts it already changed stay changed. |
Everything in your bucket is encrypted before it leaves the server. Rowsafe's servers never see your data or your passphrase.
Docker
Use the ClickHouse agent image next to the official clickhouse/clickhouse-server image. It is built on the same image, so the clickhouse program that Proof and Rewind copies run is ClickHouse's own, of your version. There is one image per version: ghcr.io/rowsafe/agent:clickhouse26.8, clickhouse26.3 and clickhouse25.8. These tags follow the newest Rowsafe release. To pin one, use the exact tag instead, like <version>-clickhouse26.8.
services:
clickhouse:
image: clickhouse/clickhouse-server:26.8
environment:
CLICKHOUSE_USER: admin
CLICKHOUSE_DEFAULT_ACCESS_MANAGEMENT: "1" # admin may create Rowsafe's user
env_file: clickhouse.env # CLICKHOUSE_PASSWORD=...
volumes:
- chdata:/var/lib/clickhouse
rowsafe-agent:
image: ghcr.io/rowsafe/agent:clickhouse26.8
hostname: db-1
environment:
ROWSAFE_CLICKHOUSE_URL: http://clickhouse:8123
ROWSAFE_CLICKHOUSE_GATEWAY_URL: http://rowsafe-agent:9010
env_file: rowsafe-agent.env
volumes:
- rowsafe-state:/var/lib/rowsafeClickHouse sends its backups to the agent on port 9010, inside the compose network. Don't publish that port. Then, once:
docker compose exec rowsafe-agent rowsafe-agent clickhouse login --port 8123 --admin-user admin
rowsafe adopt events --host db-1 --engine clickhouse --port 8123
rowsafe apply eventsclickhouse login reads the administrator's password from what you type (one line). It is used once and never saved. The full example is deploy/docker/compose.clickhouse.example.yml.
Update the agent
The agent never updates itself in Docker. To update it, download the newest image and recreate only the agent's container. ClickHouse keeps running:
docker compose pull rowsafe-agent && docker compose up -d rowsafe-agentName the service: a plain docker compose pull also downloads a newer clickhouse-server image if there is one, and docker compose up -d then restarts ClickHouse to use it. The server's page in the dashboard says when a new version is out.
To get Update now in the dashboard instead, add the container control service. It holds the Docker socket, so the agent never gets it:
services:
rowsafe-docker-control:
image: ghcr.io/rowsafe/docker-control:latest
restart: unless-stopped
depends_on: [clickhouse]
environment:
ROWSAFE_CONTROL_SERVICE: clickhouse # your ClickHouse service's name
ROWSAFE_CONTROL_ALLOW_UIDS: "101" # the agent image's clickhouse user
ROWSAFE_CONTROL_UPDATE_ONLY: "1" # only update the agent: never stop or restart ClickHouse
volumes:
- /var/run/docker.sock:/var/run/docker.sock
- rowsafe-control:/run/rowsafe-control
network_mode: none
read_only: true
cap_drop: [ALL]
security_opt: ["no-new-privileges:true"]
rowsafe-agent:
volumes:
- rowsafe-state:/var/lib/rowsafe
- rowsafe-control:/run/rowsafe-control # added
volumes:
rowsafe-control:When an owner or admin clicks Update now and confirms, the control service checks the new release's list of agent images against Rowsafe's release signature, downloads ghcr.io/rowsafe/agent for clickhouse26.8 at the digest in that list, and recreates only the agent's container with the same settings. If the new agent doesn't report in within five minutes, the old container is started again and the dashboard says why. ClickHouse keeps running throughout. A backup or Proof that was running is reported as failed and runs again.
With ROWSAFE_CONTROL_UPDATE_ONLY, the control service can only update the agent's container: it refuses to stop, start or restart ClickHouse, whoever asks. How it works and what it can and can't do: Update the agent from the dashboard and Threat model.
Restore without Rowsafe
The agent is open source, and your backups are ClickHouse's own backup format once decrypted. In the bucket, every file and its name are encrypted, so use the agent to download a backup: it decrypts both. On any Linux server with the agent and your bucket's settings: the ROWSAFE_REPO_* lines of /etc/rowsafe/agent.env, or the same values set by hand on a new server (the passphrase is the one you saved at install):
set -a; . /etc/rowsafe/agent.env; set +a
# --stanza is the database's "Bucket folder" (on its Settings page in the dashboard).
rowsafe-agent clickhouse download-backup --stanza analytics
rowsafe-agent clickhouse download-backup --stanza analytics --label 20261002-010000F_20261002-140000D --to /var/lib/clickhouse/backups/restore
chown -R clickhouse:clickhouse /var/lib/clickhouse/backups/restoreThe first command lists the backups and Marks. The second decrypts one, and for a differential backup (or a Mark) the full backup it builds on too, then prints the RESTORE statement to run. With /var/lib/clickhouse/backups/ in ClickHouse's backups.allowed_path, run it in clickhouse-client, for example:
RESTORE ALL FROM File('/var/lib/clickhouse/backups/restore/20261002-010000F_20261002-140000D/')
SETTINGS base_backup = File('/var/lib/clickhouse/backups/restore/20261002-010000F/')Limits
- No restore to any second. Restores go to a backup or a Mark (see above).
- Clusters: Rowsafe protects the ClickHouse server where the agent runs. A cluster with several shards isn't backed up as one yet. Replicated tables are backed up from this server's replica.
- Bringing rows back isn't one transaction. If it stops half-way, what was written stays; running it again finishes the job without duplicates.
- Rows aren't brought back into SummingMergeTree, AggregatingMergeTree, CollapsingMergeTree or VersionedCollapsingMergeTree tables: there, adding a row changes totals or cancels other rows. Restore a copy and copy what you need yourself.
- Tables that read from another system (Kafka, MySQL, PostgreSQL, S3, URL and similar) and the views fed by them are backed up but left out of Proof and Rewind copies, so a test never reads from or writes to that system.
- When a backup has replicated tables, Proof and Rewind copies start their own ClickHouse Keeper for them. Its internal port listens on all of the server's addresses while the copy exists (ClickHouse has no setting to limit it); the task log says so. Backups without replicated tables don't start one.
- ClickHouse 24.8 or newer.
- Rewinding the whole database in place, restarts from the dashboard, standbys, pooling, updates and Logs are PostgreSQL-only for now.
Protect MongoDB
Backups to your own bucket, restore to any second, weekly Proof, Rewind copies and Pulse for MongoDB 6.0, 7.0 and 8.0 on your own server or in Docker.
Move in from a managed database
Move a PostgreSQL database from DigitalOcean, RDS, Aurora, Supabase, Neon, Heroku, Render, Railway, Crunchy Bridge, Cloud SQL or Azure onto your own server, with near-zero downtime.