You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Docker中Scylla资源利用率低、写入性能差及超时问题求助

Docker环境下Scylla写入性能瓶颈排查与优化方案

问题概述

在Docker环境中对单节点Scylla进行压测时,每秒成功写入请求无法突破2000次,超出阈值后触发超时错误,即便分配了大量资源,Scylla也未充分利用。相同客户端向Cassandra压测时,可实现40000次/秒的成功写入。

超时错误信息:

Operation timed out for keyspace1.messagedb - received only 0 responses from 1 CL=LOCAL_SERIAL.

当前配置详情

Docker Compose配置片段

some-scylla:
  image: scylladb/scylla:latest
  container_name: some-scylla
  restart: always
  ports:
    - "9333:10000"
  command: [
    "--smp", "16",
    "--memory", "10G",
    "--experimental", "0",
    "--io-setup", "0",
    "--disable-version-check"
  ]
  volumes:
    - scylla3:/var/lib/scylla
    - ./scylla.d/memory.conf:/etc/scylla.d/memory.conf
    - ./scylla.d/io.conf:/etc/scylla.d/io.conf
    - ./scylla.yaml:/etc/scylla/scylla.yaml
    - ./scylla.d/io_properties.yaml:/etc/scylla.d/io_properties.yaml
    - ./scylla-server:/etc/default/scylla-server

memory.conf配置

MEM_CONF="--memory=10G"

cpuset.conf配置

CPUSET="--smp 16"

Scylla.yaml配置

num_tokens: 1024
commitlog_sync: periodic
commitlog_sync_period_in_ms: 200
commitlog_segment_size_in_mb: 1024
schema_commitlog_segment_size_in_mb: 4096
seed_provider:
    - class_name: org.apache.cassandra.locator.SimpleSeedProvider
      parameters:
          - seeds: "127.0.0.1"
listen_address: localhost
native_transport_port: 9042
native_shard_aware_transport_port: 19042
read_request_timeout_in_ms: 2000
write_request_timeout_in_ms: 2000
cas_contention_timeout_in_ms: 1000
endpoint_snitch: SimpleSnitch
rpc_address: localhost
rpc_port: 9160
api_port: 10000
api_address: 127.0.0.1
batch_size_warn_threshold_in_kb: 1024
batch_size_fail_threshold_in_kb: 8192
partitioner: org.apache.cassandra.dht.Murmur3Partitioner
commitlog_total_space_in_mb: -1
developer_mode: false
murmur3_partitioner_ignore_msb_bits: 12
force_schema_commit_log: true
task_ttl_in_seconds: 10
consistent_cluster_management: true
strict_is_not_null_in_views: true
api_ui_dir: /opt/scylladb/swagger-ui/dist/
api_doc_dir: /opt/scylladb/api/api-doc/

客户端(gocqlx)配置

cluster := gocql.NewCluster(hosts...)
cluster.Keyspace = Keyspace
cluster.Consistency = gocql.One
cluster.ProtoVersion = 4
cluster.Port = port
cluster.RetryPolicy = &gocql.SimpleRetryPolicy{NumRetries: 1}
cluster.PoolConfig.HostSelectionPolicy = gocql.HostPoolHostPolicy(
    hostpool.NewEpsilonGreedy(nil, 0, &hostpool.LinearEpsilonValueCalculator{}),
)
cluster.NumConns = 64
cluster.MaxRoutingKeyInfo = 1000

注意:gocql每个连接固定发送2个请求,此设置无法修改,测试仅使用单节点。

优化方案

1. 调整Docker资源隔离配置

Scylla对CPU、IO的隔离敏感,默认Docker配置可能限制性能:

  • 添加cpus: "16"明确分配CPU核心,避免调度竞争
  • 开启privileged: true或配置devices映射,确保Scylla能直接访问存储设备的IO队列
  • 添加mem_swappiness: 0禁用内存交换,防止性能骤降

2. 修正Scylla的IO配置

--io-setup=0会跳过自动IO优化,需调整:

  • 启用--io-setup=1让Scylla自动检测并配置IO参数,替代手动生成的io.conf和io_properties.yaml
  • 若必须手动配置,确保io_properties.yaml设置正确的队列深度、IO调度器(推荐noop或mq-deadline)

3. 优化Scylla.yaml关键参数

  • 调整超时时间:当前write_request_timeout_in_ms=2000可能过短,压测时可临时调至5000,长期需结合实际场景调整
  • 启用分片感知传输:客户端连接native_shard_aware_transport_port:19042,避免跨分片请求的性能损耗
  • 调整commitlog参数:将commitlog_sync_period_in_ms调至1000减少磁盘写入频率;commitlog_total_space_in_mb设为2048(内存的20%左右),避免commitlog无限增长
  • 修改num_tokens:单节点场景下num_tokens=1024过大,建议改为256,减少分片管理开销

4. 优化客户端连接配置

  • 由于gocql每个连接发2个请求,当前NumConns=64对应128个并发请求,远低于Scylla处理能力,可将NumConns调至256或更高(建议按CPU核心数*4配置)
  • 启用cluster.PoolConfig.PerHostPoolSize并设置为32,增加每个节点的连接池大小
  • 升级gocql版本至最新,修复旧版本与Scylla的兼容性问题

5. 验证资源利用率

  • 用docker stats some-scylla查看CPU、内存、磁盘IO使用率,确认是否存在资源瓶颈
  • 进入容器执行nodetool status、nodetool tpstats查看Scylla内部线程池状态,排查任务堆积情况
  • 用iostat -x 1检查磁盘IO延迟,若%util接近100%或await过高,需更换更快的存储(如NVMe)

内容的提问来源于stack exchange,提问作者Mohammad Hadi Setak

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.24 21:00:58