Milvus ID与文件名不匹配导致图像检索失败求助
问题:Milvus返回ID与本地文件名不匹配导致图像检索失败
运行图像检索程序时,出现错误提示:
Image with ID 451700669611478586 not found in the filenames list.
Milvus返回的ID无法对应本地目录中的文件名,导致无法加载检索到的图像。
原因分析
- 你的Milvus集合中
_id字段设置了auto_id=True,Milvus会自动生成全局唯一的大整数主键,而非你插入时传入的自定义img_id。 - 插入数据时你传入的
ids列表实际存入了image_id字段,但检索时你取的是result.id(Milvus自动生成的主键),用这个ID去索引本地filenames列表自然不匹配。 filenames列表仅在程序运行时存在,重启后丢失,无法持久化关联Milvus中的数据。
解决方案
最可靠的方式是将文件名直接存储到Milvus集合中,检索时直接返回文件名,彻底避免ID映射问题。以下是修正后的完整代码:
修正后的代码
import os import numpy as np import cv2 import matplotlib.pyplot as plt from pymilvus import Collection, CollectionSchema, FieldSchema, DataType, connections, utility from pymilvus.exceptions import IndexNotExistException from tensorflow.keras.applications import ResNet50 from tensorflow.keras.applications.resnet50 import preprocess_input # Step 1: Load the ResNet50 Model model = ResNet50(weights='imagenet', include_top=False, pooling='avg') # Step 2: Connect to Milvus connections.connect("default", host='localhost', port='19530') # Step 3: Define the Milvus Collection Schema - 新增filename字段存储文件名 fields = [ FieldSchema(name="_id", dtype=DataType.INT64, is_primary=True, auto_id=True), FieldSchema(name="embedding", dtype=DataType.FLOAT_VECTOR, dim=2048), FieldSchema(name="filename", dtype=DataType.VARCHAR, max_length=255) # 存储图像文件名 ] schema = CollectionSchema(fields, description="Facial embeddings collection") collection_name = "facial_embeddings" # Create the collection if it doesn't exist if collection_name not in utility.list_collections(): collection = Collection(name=collection_name, schema=schema) else: collection = Collection(name=collection_name) # Step 4: Function to Extract Features def extract_features(image_path): img = cv2.imread(image_path) img = cv2.resize(img, (224, 224)) img = np.expand_dims(img, axis=0) img = preprocess_input(img) features = model.predict(img) return features.flatten() # Step 5: Insert Features into Milvus - 调整插入逻辑,存储文件名 def insert_features(image_folder): embeddings = [] filenames = [] print("Starting feature extraction...") total_files = len([f for f in os.listdir(image_folder) if f.endswith(('.jpg', '.png'))]) count = 0 for img_file in os.listdir(image_folder): if img_file.endswith(('.jpg', '.png')): count += 1 img_path = os.path.join(image_folder, img_file) features = extract_features(img_path) embeddings.append(features) filenames.append(img_file) print(f"Extracted features from {img_file} ({count}/{total_files})") # 插入数据:对应字段顺序为embedding, filename print("Inserting features into Milvus...") collection.insert([embeddings, filenames]) collection.flush() # 确保数据写入磁盘 print("Insertion complete.") # Create an index for the embedding field if it doesn't exist try: collection.index() print("Index already exists.") except IndexNotExistException: index_params = { "index_type": "IVF_FLAT", "metric_type": "L2", "params": {"nlist": 100} } collection.create_index(field_name="embedding", index_params=index_params) print("Index created.") collection.load() print("Collection loaded.") # Step 6: Find Similar Images - 指定返回filename字段 def find_similar_images(user_image_path, top_k=3): collection.load() user_features = extract_features(user_image_path) search_params = { "metric_type": "L2", "params": {"nprobe": 10} } results = collection.search( data=[user_features], anns_field="embedding", param=search_params, limit=top_k, output_fields=["filename"] # 检索时返回文件名 ) return results # Step 7: Display Images - 直接使用Milvus返回的文件名 def display_similar_images(user_image_path, similar_results, image_folder): user_img = cv2.imread(user_image_path) if user_img is None: print(f"Failed to load user image: {user_image_path}") return user_img = cv2.cvtColor(user_img, cv2.COLOR_BGR2RGB) plt.figure(figsize=(10, 5)) plt.subplot(2, 4, 1) plt.imshow(user_img) plt.title("User Image") plt.axis('off') for i, result in enumerate(similar_results[0]): filename = result.entity.get("filename") img_path = os.path.join(image_folder, filename) similar_img = cv2.imread(img_path) if similar_img is not None: similar_img = cv2.cvtColor(similar_img, cv2.COLOR_BGR2RGB) plt.subplot(2, 4, i + 2) plt.imshow(similar_img) plt.title(f"Match {i + 1}\nDistance: {result.distance:.4f}") plt.axis('off') else: print(f"Image not found: {img_path}") plt.tight_layout() plt.show() # Step 8: Main Execution if __name__ == "__main__": image_folder = r'C:\Users\hp\Desktop\temp-folder\test_folder' insert_features(image_folder) user_image_path = r'C:\Users\hp\Desktop\temp-folder\extracted_image_12.jpg' similar_results = find_similar_images(user_image_path) display_similar_images(user_image_path, similar_results, image_folder)
关键修改点说明
- Schema调整:新增
filename字段,将图像文件名直接存储到Milvus中,持久化关联关系。 - 插入逻辑:移除自定义
ids列表,插入时将文件名传入Milvus的filename字段。 - 检索逻辑:通过
output_fields参数指定返回filename,无需再依赖本地filenames列表。 - 显示逻辑:直接从检索结果中获取文件名,构建图像路径加载图像,彻底解决ID不匹配问题。
内容的提问来源于stack exchange,提问作者lazy panda
相关产品推荐
相关产品推荐

