Restore a rocks database from object storage
readme.md

eat-rocks #

docs.rs

restore a rocks backup from s3-compatible object storage.

rocks has built-in backup/restore, but it expects a local filesystem. mounting object storage to a filesystem works, but it's really annoying.

eat-rocks talks to object storage directly (and with high default concurrency) so it can be pretty fast at getting your data back.

eat-rocks can also push rocks BackupEngine-compatible backups to object storage from a checkpoint folder.

cli #

# restore latest from a public bucket (subdomain style)
eat-rocks --endpoint https://constellation.t3.storage.dev restore /data/rocksdb

# list available backups
eat-rocks --endpoint https://constellation.t3.storage.dev list

# restore a specific backup
eat-rocks --endpoint https://constellation.t3.storage.dev restore \
  --backup-id 3 /data/rocksdb

# authenticated access (or set AWS_ACCESS_KEY_ID / AWS_SECRET_ACCESS_KEY)
eat-rocks \
  --endpoint https://constellation.t3.storage.dev \
  --access-key-id AKIA... \
  --secret-access-key wJal... \
  restore /data/rocksdb

# path-style (minio, localstack)
eat-rocks \
  --endpoint http://localhost:9000 \
  --bucket mybucket \
  restore /data/rocksdb

# limit concurrency (poor connection, etc)
eat-rocks --endpoint https://constellation.t3.storage.dev \
  restore --concurrency 8 /data/rocksdb

# always overwrite SSTs at `target`, even if they exist and crc32c matches
eat-rocks --endpoint https://constellation.t3.storage.dev \
  restore --always-download /data/rocksdb

# push a new backup from checkpoint, retaining the newest 5
# (you have to run rocks' `Checkpoint::CreateCheckpoint` yourself first)
eat-rocks --endpoint https://constellation.t3.storage.dev \
  push --retain 5 /data/checkpoints/2026-08-02

# delete each checkpoint file as it's pushed, then remove the directory.
# checkpoint files are hard links, so this keeps peak disk usage down.
eat-rocks --endpoint https://constellation.t3.storage.dev \
  push --consume-checkpoint /data/checkpoints/2026-08-02

# manage retention
eat-rocks --endpoint https://constellation.t3.storage.dev retain 5
eat-rocks --endpoint https://constellation.t3.storage.dev delete 3
eat-rocks --endpoint https://constellation.t3.storage.dev cleanup --dry-run

lib #

use eat_rocks::{public_bucket, restore};

let store = public_bucket("https://constellation.t3.storage.dev")?;
restore(store, "", "/data/rocksdb".as_ref(), Default::default()).await?;

or bring your own ObjectStore implementation (S3, GCS, Azure, local filesystem, ...):

let store: Arc<dyn ObjectStore> = /* store: all you */;
eat_rocks::restore(store, "", "/data/rocksdb".as_ref(), Default::default()).await?;

to push a backup, you first have to ask rocks to create a checkpoint. the checkpoint directory must be on the same filesystem as the database itself to avoid full file copying.

rocksdb::checkpoint::Checkpoint::new(&db)?.create_checkpoint(&cp_dir)?;
eat_rocks::push(store, "", &cp_dir, eat_rocks::PushOptions {
    retain: std::num::NonZeroUsize::new(5),
    consume_checkpoint: true,
    ..Default::default()
}).await?;

features #

  • cli: enable deps to build the binary
  • easy (default): public_bucket() convenience function with aws store backend

differences #

  • incremental by default: checks if an existing local SST file already exists and crc32c matches, to avoid redundant downloads. pass --always-download to force overwriting (config: always_download: true).

  • strict replace/create behaviour: refuses to clobber an existing rocksdb database unless you pass --replace, and refuses to create a new database if you do pass --replace. pass --replace=force to create-or-replace a database.

  • cleanup of orphaned file: when replacing an existing rocksdb database, eat-rocks will delete old no-longer-used rocksdb-like files from the target directory. rocks' own restore function leaves them in place and they never get cleaned up.

TODO: keep an instance in memory to avoid reading all metas from bucket on every push (like rocks)

be aware #

  • running a push concurrently with a cleanup can race, leaving the racing push in an invalid state. don't run them at the same time if you can avoid it.

  • also and even further on the edge: if you concurrently run a cleanup with very short stale_marker_after or a very, very long-running push, lasting more than stale_marker_after... then the cleanup could delete files of that ongoing push. by default that would be a push taking over 7 days, and doing a concurrent cleanup. avoid both if you can (you can).

  • if a meta file gets corrupted, pushes and cleanups will fail. this isn't essential for push, we could probably skip and proceed (risking unlimited storage growth?), but it is essential for cleanup to avoid deleting files that the corrupted meta references. open to suggested changes. note that corruption of files in object storage is extremely rare.

  • if two metas claim different crc32cs for the same file, eat-rocks will currently just pick one, non-deterministically (last fetch completion wins). this shouldn't happen, but eat-rocks should probably error if it does. open to a PR fixing this.

changes #

see ./changelog.md

hacking #

see ./hacking.md

cli build #

the cli feature flag is required to build the cli

cargo build --release --features cli

license #

Dual-licensed under MIT and Apache 2.0.

SPDX-License-Identifier: MIT OR Apache-2.0