Milvus在AWS EKS集群无法利用多节点内存加载集合问题排查
问题
我有1000万个维度为1536的embedding,计算后内存需求小于128GB。在AWS EKS集群上运行Milvus,集群包含2台m6i.4xlarge节点(每台64GB内存),理论内存充足。
我理解集合需要分区才能加载到多节点,因此设置了分区键,代码如下:
fields = [ FieldSchema(name="id", dtype=DataType.VARCHAR, max_length=24, is_primary=True), FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=VECTOR_DIMENSION), FieldSchema(name="document", dtype=DataType.VARCHAR, max_length=30000), ] schema = CollectionSchema( fields, description="Vector store collection for arxiv", partition_key_field="document", n_partitions=16 ) milvus_collection = Collection(name=MILVUS_COLLECTION_NAME, schema=schema)
添加数据并创建索引:
index_params = { "index_type": "IVF_FLAT", "metric_type": "L2", "params": {"nlist": 1024} } milvus_collection.create_index(field_name="embedding", index_params=index_params)
执行加载命令:
milvus_collection.load()
加载到38%时出现错误:
pymilvus.exceptions.MilvusException: <MilvusException: (code=65535,
message=show collection failed: load segment failed, OOM if load,
maxSegmentSize = 1576.6993761062622 MB, memUsage = 55411.66015625 MB,
predictMemUsage = 56988.35953235626 MB, totalMem = 63273.46875 MB
thresholdFactor = 0.900000)>
请问:
- 为何只能使用单节点的内存?
- 是否有方法查看每个分区的存储位置和大小?
解决方案
1. 单节点内存受限的原因
- Segment无法跨节点存储:Milvus的单个Segment只能被分配到单个节点存储,即使设置了分区键,若某个分区对应的Segment过大(错误信息中单个Segment接近1.5GB),加载时会占用单节点大量内存,触发OOM阈值。
- 额外内存开销未被计算:仅计算向量内存是不够的,IVF_FLAT索引的nlist=1024会带来额外内存占用,加上
document字段的大VARCHAR元数据,实际内存需求远超单纯的向量内存,导致单节点负载接近90%阈值(错误信息中thresholdFactor=0.9)。 - 数据分布不均匀:
n_partitions=16不代表数据会均匀分配,若document字段的哈希分布失衡,部分分区数据量远大于其他分区,会导致单个节点承载过多Segment,内存耗尽。
2. 查看分区存储位置和大小的方法
- 查询分区数据量:使用Python SDK的
describe_partitions方法获取分区基本信息:partitions = milvus_collection.describe_partitions() for p in partitions: print(f"分区名称: {p.name}, 数据量: {p.num_rows}") - 查看Segment节点分布:通过Milvus内置的
system.segments集合查询每个Segment的节点位置和大小:from pymilvus import Collection system_segments = Collection("system.segments") res = system_segments.query( expr=f"collection_name == '{MILVUS_COLLECTION_NAME}'", output_fields=["segment_id", "partition_name", "node_id", "size"] ) for seg in res: print(f"分区: {seg['partition_name']}, Segment ID: {seg['segment_id']}, 节点ID: {seg['node_id']}, 大小(字节): {seg['size']}") - 控制台可视化查看:若使用Milvus集群的控制台界面,可在集合详情页直接查看分区、Segment的分布情况,直观看到各节点的负载状态。
内容的提问来源于stack exchange,提问作者Qi Xiang
相关产品推荐
相关产品推荐

