You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Milvus在AWS EKS集群无法利用多节点内存加载集合问题排查

问题

我有1000万个维度为1536的embedding,计算后内存需求小于128GB。在AWS EKS集群上运行Milvus,集群包含2台m6i.4xlarge节点(每台64GB内存),理论内存充足。

我理解集合需要分区才能加载到多节点,因此设置了分区键,代码如下:

fields = [
            FieldSchema(name="id", dtype=DataType.VARCHAR, max_length=24, is_primary=True),
            FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=VECTOR_DIMENSION),
            FieldSchema(name="document", dtype=DataType.VARCHAR, max_length=30000),
        ]

schema = CollectionSchema(
        fields,
        description="Vector store collection for arxiv",
        partition_key_field="document",
        n_partitions=16
)

milvus_collection = Collection(name=MILVUS_COLLECTION_NAME, schema=schema)

添加数据并创建索引:

index_params = {
    "index_type": "IVF_FLAT",
    "metric_type": "L2",
    "params": {"nlist": 1024}
}
milvus_collection.create_index(field_name="embedding", index_params=index_params)

执行加载命令:

milvus_collection.load()

加载到38%时出现错误:

pymilvus.exceptions.MilvusException: <MilvusException: (code=65535,
message=show collection failed: load segment failed, OOM if load,
maxSegmentSize = 1576.6993761062622 MB, memUsage = 55411.66015625 MB,
predictMemUsage = 56988.35953235626 MB, totalMem = 63273.46875 MB
thresholdFactor = 0.900000)>

请问:

  1. 为何只能使用单节点的内存?
  2. 是否有方法查看每个分区的存储位置和大小?

解决方案

1. 单节点内存受限的原因

  • Segment无法跨节点存储:Milvus的单个Segment只能被分配到单个节点存储,即使设置了分区键,若某个分区对应的Segment过大(错误信息中单个Segment接近1.5GB),加载时会占用单节点大量内存,触发OOM阈值。
  • 额外内存开销未被计算:仅计算向量内存是不够的,IVF_FLAT索引的nlist=1024会带来额外内存占用,加上document字段的大VARCHAR元数据,实际内存需求远超单纯的向量内存,导致单节点负载接近90%阈值(错误信息中thresholdFactor=0.9)。
  • 数据分布不均匀:n_partitions=16不代表数据会均匀分配,若document字段的哈希分布失衡,部分分区数据量远大于其他分区,会导致单个节点承载过多Segment,内存耗尽。

2. 查看分区存储位置和大小的方法

  • 查询分区数据量:使用Python SDK的describe_partitions方法获取分区基本信息:
    partitions = milvus_collection.describe_partitions()
    for p in partitions:
        print(f"分区名称: {p.name}, 数据量: {p.num_rows}")
    
  • 查看Segment节点分布:通过Milvus内置的system.segments集合查询每个Segment的节点位置和大小:
    from pymilvus import Collection
    system_segments = Collection("system.segments")
    res = system_segments.query(
        expr=f"collection_name == '{MILVUS_COLLECTION_NAME}'", 
        output_fields=["segment_id", "partition_name", "node_id", "size"]
    )
    for seg in res:
        print(f"分区: {seg['partition_name']}, Segment ID: {seg['segment_id']}, 节点ID: {seg['node_id']}, 大小(字节): {seg['size']}")
    
  • 控制台可视化查看:若使用Milvus集群的控制台界面,可在集合详情页直接查看分区、Segment的分布情况,直观看到各节点的负载状态。

内容的提问来源于stack exchange,提问作者Qi Xiang

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 15:04:51