Cassandra Stress YAML配置文件参数解析及疑问咨询
Cassandra-stress YAML配置参数疑问解答
问题与解答
1. 分区数与总行数的计算逻辑
问题:columnspec中,分区列host、bucket_time、service的population参数分别为uniform(1..600)、uniform(1..288)、uniform(1000..2000),是否意味着压力测试完成后最大分区数为600*288*2000,执行select count(*)得到的最大行数为600*288*2000*15?
解答:
- 最大分区数:这三个字段是复合分区键,理论上的最大可能分区数确实是
600*288*2000,但实际测试中因为uniform是随机生成值,不一定能覆盖所有组合,所以实际分区数会小于等于这个最大值。 - 最大总行数:聚类列
time的cluster: fixed(15)配置表示每个分区下会固定生成15个不同的时间值,对应15行数据。因此理论上最大总行数是最大分区数乘以15,即600*288*2000*15,同样实际行数取决于测试是否覆盖了所有可能的分区组合。
2. insert.partitions: fixed(1)的含义
问题:insert配置中的partitions: fixed(1),是否表示每次插入操作仅更新1个分区?
解答:是的。partitions: fixed(1)指定了每次插入批次(batch)仅针对1个分区进行操作。结合配置中的batchtype: UNLOGGED,这个设置符合Cassandra的性能最佳实践——单分区的unlogged batch执行效率最高。
3. insert.select: fixed(10)/10的作用
问题:insert配置中的select: fixed(10)/10是什么含义?初始表为空时该参数如何运作,是否表示选取批次中100%的数据进行插入?
解答:
- 这个参数的格式为
fixed(X)/Y,其中X是每次要生成的行数,Y是用来计算插入比例的分母。fixed(10)/10表示每次插入操作会生成10行数据,并且这10行全部会被插入(比例为10/10=100%)。 - 当表为空时,cassandra-stress会直接生成指定的新数据执行插入,不会跳过任何行——因为这个参数在这里的作用是控制生成数据中实际插入的比例,分母10意味着所有生成的10行都会被插入。
附完整配置文件
# Keyspace name and create CQL # keyspace: stressexample keyspace_definition: | CREATE KEYSPACE stressexample WITH replication = {'class': 'NetworkTopologyStrategy', 'AWS_VPC_US_WEST_2': '2'}; # # Table name and create CQL # table: eventsrawtest table_definition: | CREATE TABLE eventsrawtest ( host text, bucket_time text, service text, time timestamp, metric double, state text, PRIMARY KEY ((host, bucket_time, service), time) ) WITH CLUSTERING ORDER BY (time DESC) # # Meta information for generating data # columnspec: - name: host size: fixed(32) #In chars, no. of chars of UUID population: uniform(1..600) # We have about 600 hosts with equal events per host - name: bucket_time size: fixed(18) population: uniform(1..288) # 288 potential buckets - name: service size: uniform(10..100) population: uniform(1000..2000) # 1000 - 2000 metrics per host - name: time cluster: fixed(15) - name: state size: fixed(4) # # Specs for insert queries # insert: partitions: fixed(1) # 1 partition per batch batchtype: UNLOGGED # use unlogged batches select: fixed(10)/10 # no chance of skipping a row when generating inserts # # Read queries to run against the schema # queries: pull-for-rollup: cql: select * from eventsrawtest where host = ? and service = ? and bucket_time = ? fields: samerow # pick selection values from same row in partition get-a-value: cql: select * from eventsrawtest where host = ? and service = ? and bucket_time = ? and time = ? fields: samerow # pick selection values from same row in partition
内容的提问来源于stack exchange,提问作者Pavel Orekhov
相关产品推荐
相关产品推荐

