Azure认知向量搜索(ACS)索引创建与向量查询报错求助
Azure认知向量搜索(ACS)索引创建与向量查询报错求助
我想基于向量查询实现一个Azure认知搜索(ACS)的课程目录检索功能,目前已经有一个名为courses_pd的Pandas DataFrame,包含content(课程内容)和embeddings(我用SentenceTransformer的all-MiniLM-L6-v2模型编码生成的向量)两列。
我已经在Azure Databricks notebook中编写了创建ACS索引和上传文档的代码,具体如下:
创建索引的代码片段
from azure.search.documents.indexes import SearchIndexClient from azure.search.documents.indexes.models import ( SimpleField, SearchFieldDataType, SearchableField, SearchField, VectorSearch, HnswAlgorithmConfiguration, VectorSearchProfile, SemanticConfiguration, SemanticPrioritizedFields, SemanticField, SemanticSearch, SearchIndex, AzureOpenAIVectorizer, AzureOpenAIVectorizerParameters ) from azure.core.credentials import AzureKeyCredential # Azure Cognitive Search setup service_endpoint = "https://yourserviceendpoint.search.windows.net" admin_key = "ABC" index_name = "courses-index" # Wrap admin_key in AzureKeyCredential credential = AzureKeyCredential(admin_key) # Create the index client with AzureKeyCredential index_client = SearchIndexClient(endpoint=service_endpoint, credential=credential) # Define the index schema fields = [ SimpleField(name="id", type="Edm.String", key=True), SimpleField(name="content", type="Edm.String"), SearchField( name="embedding", type=SearchFieldDataType.Collection(SearchFieldDataType.Single), searchable=True, vector_search_dimensions=384, vector_search_profile_name="myHnswProfile" ) # SearchField(name="embedding", type='Collection(Edm.Single)', searchable=True) ] # Configure the vector search configuration vector_search = VectorSearch( algorithms=[ HnswAlgorithmConfiguration( name="myHnsw" ) ], profiles=[ VectorSearchProfile( name="myHnswProfile", algorithm_configuration_name="myHnsw" ) ] ) # Create the index index = SearchIndex( name=index_name, fields=fields, vector_search=vector_search ) # Send the index creation request index_client.create_index(index) print(f"Index '{index_name}' created successfully.")
上传文档的代码片段
from azure.search.documents import SearchClient # Generate embeddings and upload data search_client = SearchClient(endpoint=service_endpoint, index_name=index_name, credential=credential) documents = [] for i, row in courses_pd.iterrows(): document = { "id": str(i), "content": row["content"], "embedding": row["embeddings"] # Ensure embeddings are a list of floats } documents.append(document) # Upload documents to the index search_client.upload_documents(documents=documents) print(f"Uploaded {len(documents)} documents to Azure Cognitive Search.")
查询时遇到的问题
现在我尝试查询时遇到了多个错误。不管是直接用原始字符串查询,还是先把查询字符串用模型编码成向量再查询,都会返回<iterator object azure.core.paging.ItemPaged at 0x7fcf9f086220>,同时日志显示查询失败。
我的查询代码如下:
from azure.search.documents.models import VectorQuery # Generate embedding for the query query = "machine learning" query_embedding = model.encode(query).tolist() # Convert to list of floats # Create a VectorQuery vector_query = VectorQuery( vector=query_embedding, k=3, # Number of nearest neighbors fields="embedding" # Name of the field where embeddings are stored ) # Perform the search results = search_client.search( vector_queries=[vector_query], select=["id", "content"] ) # Print the results for result in results: print(f"ID: {result['id']}, Content: {result['content']}")
具体错误信息
运行上述查询代码后,我得到了以下错误:
vector is not a known attribute of class <class 'azure.search.documents._generated.models._models_py3.VectorQuery'> and will be ignored k is not a known attribute of class <class 'azure.search.documents._generated.models._models_py3.VectorQuery'> and will be ignored HttpResponseError: (InvalidRequestParameter) The vector query's 'kind' parameter is not set.
我尝试在VectorQuery中添加kind = 'vector'参数,但仍然报错说kind未设置!
我已经在Azure门户中确认索引和文档都已成功创建,索引结构显示正常(能看到定义的id、content和embedding字段)。
我应该是在索引创建或者查询的方式上出了问题,查了官方文档和GitHub代码库都没找到相关的解决方法,我是刚接触这个技术,希望能得到社区的帮助,谢谢!
备注:内容来源于stack exchange,提问作者Strayhorn
相关产品推荐
相关产品推荐

