Cassandra能否像MongoDB的Simple replication模式那样仅复制不分片?
Hey there! Let's break down your two questions about Cassandra vs MongoDB deployment and indexing, since I've worked with both quite a bit.
Short answer: Yes, but it's not Cassandra's "native" sweet spot—and it works a bit differently than MongoDB's master-slave setup.
MongoDB's Simple Replication is a single-master, multiple-slave setup where all writes go to the master, and replicas sync the full dataset. For Cassandra to mimic this "all data on every node" behavior, you can do the following:
- Create a keyspace with a replication factor equal to the total number of nodes in your cluster, using
SimpleStrategy(for a single data center). For example:
If you have 3 nodes, every partition (row) will be stored on all 3 nodes. This means there's no "sharding" in the sense of data being split across nodes—each node holds the full dataset.CREATE KEYSPACE my_keyspace WITH replication = { 'class': 'SimpleStrategy', 'replication_factor': 3 }; - Unlike MongoDB's single-master model, Cassandra is decentralized: every node can accept writes and reads (as long as your consistency level allows it). There's no single point of failure here, which is a big difference from MongoDB's master-slave.
A quick caveat: This setup negates one of Cassandra's biggest strengths—horizontal scalability. If you're running a replication-only cluster, adding more nodes won't help with storage or throughput, since every node still holds all data. It's only useful for extreme high availability, not scaling.
Yes, that's fundamental to how Cassandra works—with a small exception for the replication-only setup I mentioned above.
Cassandra uses a token ring to distribute data: every partition key is hashed to a token, and each node in the cluster is responsible for a range of tokens. When you have more than one node, the token ring splits into ranges, and data is spread across nodes based on those token ranges.
Why is this mandatory? Because Cassandra is designed from the ground up to be a distributed, scalable database. Without partitioning by partition key, you'd have to store all data on every node (the replication-only setup), which doesn't scale. The partition key is how Cassandra knows which node(s) hold a given piece of data—this is what enables low-latency reads/writes by targeting specific nodes instead of scanning the entire cluster.
The only exception is the replication-only setup I talked about earlier, but that's not a standard deployment for Cassandra. For any production cluster where you're using multiple nodes to scale storage or throughput, data will always be split across nodes using partition keys.
Let's tie this back to your secondary index question, since it's closely linked to partitioning:
- MongoDB: Secondary indexes are stored on each node. In a sharded cluster, querying a secondary index without a shard key triggers a scatter-gather operation—sending the query to all shards and aggregating results. In a non-sharded (simple replication) setup, the master holds the authoritative index, and replicas sync it.
- Cassandra: Secondary indexes are local to each node. Each node only indexes the data it holds (its token range). If you run a secondary index query without specifying a partition key, Cassandra has to query every node in the cluster (a "full cluster scan") to collect results, which is extremely slow and not recommended for production.
This is why Cassandra strongly recommends using partition keys for all frequent queries, and suggests alternatives like materialized views or denormalization instead of relying on secondary indexes for cross-node queries. MongoDB's secondary indexes are more flexible for ad-hoc queries, but they can become a bottleneck in large sharded clusters due to scatter-gather overhead.
内容的提问来源于stack exchange,提问作者emilly

