You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Milvus ID与文件名不匹配导致图像检索失败求助

问题:Milvus返回ID与本地文件名不匹配导致图像检索失败

运行图像检索程序时,出现错误提示:

Image with ID 451700669611478586 not found in the filenames list.

Milvus返回的ID无法对应本地目录中的文件名,导致无法加载检索到的图像。

原因分析

  1. 你的Milvus集合中_id字段设置了auto_id=True,Milvus会自动生成全局唯一的大整数主键,而非你插入时传入的自定义img_id。
  2. 插入数据时你传入的ids列表实际存入了image_id字段,但检索时你取的是result.id(Milvus自动生成的主键),用这个ID去索引本地filenames列表自然不匹配。
  3. filenames列表仅在程序运行时存在,重启后丢失,无法持久化关联Milvus中的数据。

解决方案

最可靠的方式是将文件名直接存储到Milvus集合中,检索时直接返回文件名,彻底避免ID映射问题。以下是修正后的完整代码:

修正后的代码

import os
import numpy as np
import cv2
import matplotlib.pyplot as plt
from pymilvus import Collection, CollectionSchema, FieldSchema, DataType, connections, utility
from pymilvus.exceptions import IndexNotExistException
from tensorflow.keras.applications import ResNet50
from tensorflow.keras.applications.resnet50 import preprocess_input

# Step 1: Load the ResNet50 Model
model = ResNet50(weights='imagenet', include_top=False, pooling='avg')

# Step 2: Connect to Milvus
connections.connect("default", host='localhost', port='19530')

# Step 3: Define the Milvus Collection Schema - 新增filename字段存储文件名
fields = [
    FieldSchema(name="_id", dtype=DataType.INT64, is_primary=True, auto_id=True),
    FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=2048),
    FieldSchema(name="filename", dtype=DataType.VARCHAR, max_length=255)  # 存储图像文件名
]
schema = CollectionSchema(fields, description="Facial embeddings collection")
collection_name = "facial_embeddings"

# Create the collection if it doesn't exist
if collection_name not in utility.list_collections():
    collection = Collection(name=collection_name, schema=schema)
else:
    collection = Collection(name=collection_name)

# Step 4: Function to Extract Features
def extract_features(image_path):
    img = cv2.imread(image_path)
    img = cv2.resize(img, (224, 224))
    img = np.expand_dims(img, axis=0)
    img = preprocess_input(img)
    features = model.predict(img)
    return features.flatten()

# Step 5: Insert Features into Milvus - 调整插入逻辑,存储文件名
def insert_features(image_folder):
    embeddings = []
    filenames = []

    print("Starting feature extraction...")
    total_files = len([f for f in os.listdir(image_folder) if f.endswith(('.jpg', '.png'))])
    count = 0
    for img_file in os.listdir(image_folder):
        if img_file.endswith(('.jpg', '.png')):
            count += 1
            img_path = os.path.join(image_folder, img_file)
            features = extract_features(img_path)
            embeddings.append(features)
            filenames.append(img_file)
            print(f"Extracted features from {img_file} ({count}/{total_files})")

    # 插入数据:对应字段顺序为embedding, filename
    print("Inserting features into Milvus...")
    collection.insert([embeddings, filenames])
    collection.flush()  # 确保数据写入磁盘
    print("Insertion complete.")

    # Create an index for the embedding field if it doesn't exist
    try:
        collection.index()
        print("Index already exists.")
    except IndexNotExistException:
        index_params = {
            "index_type": "IVF_FLAT",
            "metric_type": "L2",
            "params": {"nlist": 100}
        }
        collection.create_index(field_name="embedding", index_params=index_params)
        print("Index created.")

    collection.load()
    print("Collection loaded.")

# Step 6: Find Similar Images - 指定返回filename字段
def find_similar_images(user_image_path, top_k=3):
    collection.load()

    user_features = extract_features(user_image_path)
    search_params = {
        "metric_type": "L2",
        "params": {"nprobe": 10}
    }
    results = collection.search(
        data=[user_features],
        anns_field="embedding",
        param=search_params,
        limit=top_k,
        output_fields=["filename"]  # 检索时返回文件名
    )
    return results

# Step 7: Display Images - 直接使用Milvus返回的文件名
def display_similar_images(user_image_path, similar_results, image_folder):
    user_img = cv2.imread(user_image_path)
    if user_img is None:
        print(f"Failed to load user image: {user_image_path}")
        return
    user_img = cv2.cvtColor(user_img, cv2.COLOR_BGR2RGB)

    plt.figure(figsize=(10, 5))
    plt.subplot(2, 4, 1)
    plt.imshow(user_img)
    plt.title("User Image")
    plt.axis('off')

    for i, result in enumerate(similar_results[0]):
        filename = result.entity.get("filename")
        img_path = os.path.join(image_folder, filename)
        similar_img = cv2.imread(img_path)
        if similar_img is not None:
            similar_img = cv2.cvtColor(similar_img, cv2.COLOR_BGR2RGB)
            plt.subplot(2, 4, i + 2)
            plt.imshow(similar_img)
            plt.title(f"Match {i + 1}\nDistance: {result.distance:.4f}")
            plt.axis('off')
        else:
            print(f"Image not found: {img_path}")

    plt.tight_layout()
    plt.show()

# Step 8: Main Execution
if __name__ == "__main__":
    image_folder = r'C:\Users\hp\Desktop\temp-folder\test_folder'
    insert_features(image_folder)

    user_image_path = r'C:\Users\hp\Desktop\temp-folder\extracted_image_12.jpg'
    similar_results = find_similar_images(user_image_path)
    display_similar_images(user_image_path, similar_results, image_folder)

关键修改点说明

  1. Schema调整:新增filename字段,将图像文件名直接存储到Milvus中,持久化关联关系。
  2. 插入逻辑:移除自定义ids列表,插入时将文件名传入Milvus的filename字段。
  3. 检索逻辑:通过output_fields参数指定返回filename,无需再依赖本地filenames列表。
  4. 显示逻辑:直接从检索结果中获取文件名,构建图像路径加载图像,彻底解决ID不匹配问题。

内容的提问来源于stack exchange,提问作者lazy panda

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.19 23:20:54