You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Astra Cassandra插入大体积text字段写入失败问题排查

Cassandra迁移Astra大字段写入失败问题

问题背景

将AWS服务器上部署的自建Cassandra集群迁移至Astra Cassandra过程中,无法向Astra Cassandra插入单条包含约200万字符、体积为1.77MB的text类型列数据,后续还需要插入体量约2000万字符的更大规模数据,需明确问题解决思路。

问题现象

  • 插入操作通过Python应用执行,使用驱动版本为cassandra-driver==3.17.0
  • 插入全量1.77MB字段时返回错误,错误栈如下:
start.sh[5625]: [2022-07-12 15:14:39,336] 
INFO in db_ops: error = Error from server: code=1500
[Replica(s) failed to execute write] 
message="Operation failed - received 0 responses and 2 failures: UNKNOWN from 0.0.0.125:7000, UNKNOWN from 0.0.0.181:7000" 
info={'consistency': 'LOCAL_QUORUM', 'required_responses': 2, 'received_responses': 0, 'failures': 2}
  • 经测试,仅插入该字段一半长度的字符时,写入操作可正常执行。

表结构对比

Astra Cassandra目标表结构

token@cqlsh> describe mykeyspace.series;

CREATE TABLE mykeyspace.series (
    type text,
    name text,
    as_of timestamp,
    data text,
    hash text,
    PRIMARY KEY ((type, name, as_of))
) WITH additional_write_policy = '99PERCENTILE'
    AND bloom_filter_fp_chance = 0.01
    AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}
    AND comment = ''
    AND compaction = {'class': 'org.apache.cassandra.db.compaction.UnifiedCompactionStrategy'}
    AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}
    AND crc_check_chance = 1.0
    AND default_time_to_live = 0
    AND gc_grace_seconds = 864000
    AND max_index_interval = 2048
    AND memtable_flush_period_in_ms = 0
    AND min_index_interval = 128
    AND read_repair = 'BLOCKING'
    AND speculative_retry = '99PERCENTILE';

AWS自建Cassandra原表结构

ansible@cqlsh> describe mykeyspace.series;

CREATE TABLE mykeyspace.series (
    type text,
    name text,
    as_of timestamp,
    data text,
    hash text,
    PRIMARY KEY ((type, name, as_of))
) WITH bloom_filter_fp_chance = 0.01
    AND caching = {'keys': 'ALL', 'rows_per_partition': 'NONE'}
    AND comment = ''
    AND compaction = {'class': 'org.apache.cassandra.db.compaction.SizeTieredCompactionStrategy', 'max_threshold': '32', 'min_threshold': '4'}
    AND compression = {'chunk_length_in_kb': '64', 'class': 'org.apache.cassandra.io.compress.LZ4Compressor'}
    AND crc_check_chance = 1.0
    AND dclocal_read_repair_chance = 0.1
    AND default_time_to_live = 0
    AND gc_grace_seconds = 864000
    AND max_index_interval = 2048
    AND memtable_flush_period_in_ms = 0
    AND min_index_interval = 128
    AND read_repair_chance = 0.0
    AND speculative_retry = '99PERCENTILE';

插入数据样例

{"type": "OP", "name": "book", "as_of": "2022-03-17", "data": [{"year": 2022, "month": 3, "day": 17, "hour": 0, "quarter": 1, "week": 11, "wk_year": 2022, "is_peak": 0, "value": 1.28056854009628e-08}, .... ], "hash": "84421b8d934b06488e1ac464bd46e83ccd2beea5eb2f9f2c52428b706a9b2a10"}

上述JSON结构的data数组包含27000个结构一致的条目,单条条目格式如下:

{"year": 2022, "month": 3, "day": 17, "hour": 0, "quarter": 1, "week": 11, "wk_year": 2022, "is_peak": 0, "value": 1.28056854009628e-08}

插入操作Python代码片段

def insert_to_table(self, table_name, **kwargs):
        try:
            ...
            elif table_name == "series":
                self.session.execute(
                    self.session.prepare("INSERT INTO series (type, name, as_of, data, hash) VALUES (?, ?, ?, ?, ?)"),
                    (
                        kwargs["type"],
                        kwargs["name"],
                        kwargs["as_of"],
                        kwargs["data"],
                        kwargs["hash"],
                    ),
                )
            return True
        except Exception as error:
            current_app.logger.error('src/db/db_ops.py insert_to_table() table_name = %s error = %s', table_name, error)
            return False

根因分析与解决思路

  • 核心原因是Astra Cassandra作为托管服务,默认单条写入请求的大小阈值为1MB,超过阈值的请求会被协调节点直接拦截,副本无法收到写入请求,最终返回0响应、副本状态UNKNOWN的错误。半量数据可正常写入,是因为其大小刚好落在1MB阈值以内。
  • 针对当前1.77MB数据、后续2000万字符(约20MB)级别的数据,可选择以下落地方案:
    • 分块存储:调整表结构,在主键中增加chunk_id分块序号字段,将大数组按固定条目数/固定大小拆分为多个小块,单块大小控制在800KB以内(预留协议头、压缩开销空间),读取时拉取同主键下所有分块,在客户端拼接还原。该方案兼容性最好,无需依赖服务端配置调整,可稳定支持TB级以下的大字段存储,是Cassandra存储大字段的标准实践。
    • 调整驱动阈值:若暂时不想修改表结构,先将cassandra-driver升级至3.25以上版本,在集群连接配置中调大max_value_size_in_mb参数至业务需要的大小(注意Astra不同实例规格的单条值硬上限最高为25MB,不可超出对应规格限制),可先将写入一致性级别临时调整为LOCAL_ONE测试写入,确认稳定后再切回LOCAL_QUORUM。该方案不适合长期存储超过10MB的单条数据,会显著升高读取延迟、增加集群compaction压力。
    • 客户端压缩存储:在Python侧先将data字段的JSON字符串用LZ4/zlib压缩为二进制,将字段类型从text修改为blob存储。常规时序类JSON数据压缩率可达70%-85%,1.77MB文本压缩后仅200-300KB,2000万字符文本压缩后约2-3MB,可适配绝大多数场景,同时能降低存储成本,读取时在客户端解压还原即可。
  • 额外注意:原自建集群使用SizeTieredCompaction策略,Astra默认使用UnifiedCompaction策略,大字段写入后如果出现读延迟升高,可联系Astra技术支持调整表的compaction参数适配大字段场景,不要自行修改compaction策略,托管服务的参数变更有平台侧校验规则,自行修改可能导致表写入异常。

内容的提问来源于stack exchange,提问作者R13mus

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.08.26 15:24:34