求助:从3节点Cassandra集群向6节点新集群恢复快照
Hey there! I see you're new to Cassandra and need to restore snapshots from a 3-node dev cluster to a brand new 6-node production cluster, copying specific keyspaces. Let's break this down into straightforward steps that even a beginner can follow:
First things first—you need valid snapshots of the keyspaces you want to copy. If you haven't taken them yet, run this command on every single dev node for each target keyspace:
nodetool snapshot -t dev_to_prod_snapshot <your_keyspace_name>
Replace <your_keyspace_name> with each keyspace you need to migrate, and dev_to_prod_snapshot can be any name you want to label this snapshot set (it'll make it easier to find later).
Once the snapshots finish up, you'll find the data files in this directory on each dev node (the exact path might vary if you customized your Cassandra config—check cassandra.yaml for data_file_directories if you're unsure):
/var/lib/cassandra/data/<your_keyspace_name>/<table_name>-<random_uuid>/snapshots/dev_to_prod_snapshot/
Important: Each dev node only holds a subset of your cluster's data, so you need to grab the snapshot directories from all 3 dev nodes for every table in your target keyspaces.
Next, copy all those snapshot directories from your dev nodes to a temporary folder on one of your production nodes (or a machine that can reach the prod cluster's CQL port). Make sure the files are owned by the cassandra user once they're transferred—this avoids permission issues later:
sudo chown -R cassandra:cassandra /path/to/temp/prod_snapshots/
Cassandra needs to know the structure of your data before you can load it, so let's export the schema from dev and apply it to prod:
- On a dev node, export the schema for each keyspace:
cqlsh -e "DESCRIBE KEYSPACE <your_keyspace_name>" > <your_keyspace_name>_schema.cql - Copy this
.cqlfile to your prod cluster, then edit it to update the replication strategy for your 6-node setup. For example, if you want 3 replicas per piece of data, change the replication section to:WITH replication = {'class': 'SimpleStrategy', 'replication_factor': 3}; - Apply the schema to prod using
cqlsh:cqlsh -f <your_keyspace_name>_schema.cql
sstableloader The sstableloader tool is the safest way to load SSTables into a new cluster—especially since your prod cluster has more nodes than dev, so token ranges will be different. It automatically streams data to the correct prod nodes and handles replication for you.
For each table's snapshot directory, run this command (replace the seed node addresses with your prod cluster's seed nodes):
sstableloader -d prod-seed-1,prod-seed-2,prod-seed-3 /path/to/temp/prod_snapshots/<your_keyspace_name>/<table_name>-<random_uuid>/snapshots/dev_to_prod_snapshot/
Repeat this for every table in every keyspace you're migrating.
Don't skip this step—you want to make sure everything worked:
- Check that all prod nodes are up and healthy with:
All nodes should shownodetool statusUN(Up/Normal) in the status column. - Run a repair on each keyspace to ensure all replicas have consistent data:
nodetool repair <your_keyspace_name> - Use
cqlshto run some sample queries (e.g.,SELECT * FROM <your_keyspace_name>.<your_table> LIMIT 10;) to confirm your data is present.
- Test this entire process in a staging environment first before touching production—no one wants unexpected issues in prod!
- Make sure your dev and prod Cassandra versions are compatible (ideally the same major version) to avoid SSTable format conflicts.
- Take snapshots during a period of low write activity on dev to avoid partial or inconsistent data in the snapshot.
内容的提问来源于stack exchange,提问作者user9420624

