如何用Python从Azure Search Index导出全部/部分文档至DataFrame/JSON?
从Azure Cognitive Search索引导出数据到DataFrame/JSON的Python方案
核心实现脚本
首先确保安装Azure Search的Python SDK:
pip install azure-search-documents azure-identity
以下是可直接复用的代码,支持导出全部文档或基于过滤条件导出,最终转为DataFrame或JSON:
from azure.search.documents import SearchClient from azure.identity import DefaultAzureCredential import pandas as pd # 初始化SearchClient service_endpoint = "你的搜索服务端点" index_name = "你的索引名称" # 两种认证方式选其一:用默认身份认证,或直接传入API密钥 credential = DefaultAzureCredential() # credential = ApiKeyCredential(api_key="你的API密钥") search_client = SearchClient(endpoint=service_endpoint, index_name=index_name, credential=credential) def export_search_documents(filter_query=None): """ 导出搜索索引文档 :param filter_query: 过滤条件(遵循OData语法),示例:"category eq 'books'",不传则导出全部 :return: 包含所有目标文档的列表 """ all_documents = [] skip = 0 page_size = 1000 # Azure Search单次请求最大返回1000条,为服务端硬限制 while True: # search="*"匹配所有文档,select="*"返回所有字段 results = search_client.search( search="*", filter=filter_query, select="*", top=page_size, skip=skip ) page_docs = list(results) if not page_docs: break all_documents.extend(page_docs) skip += page_size return all_documents # 示例1:导出全部文档并转为DataFrame all_docs = export_search_documents() df = pd.DataFrame(all_docs) print(df.head()) # 示例2:按过滤条件导出并保存为JSON filtered_docs = export_search_documents(filter_query="price lt 50 and category eq 'electronics'") import json with open("filtered_docs.json", "w", encoding="utf-8") as f: json.dump(filtered_docs, f, ensure_ascii=False, indent=2)
你关心的疑问解答
page_size的1000限制是硬限制吗?
是的,Azure Cognitive Search的Search API单次请求最多返回1000条结果,这是服务端设定的上限,无法突破。所以当索引中文档数量超过1000时,必须通过skip参数分页循环获取。page_size和Document的区别?
page_size(对应代码中的top参数)是单次API请求返回的文档数量上限;- Document就是你所说的索引中的"行项",每个Document是索引中的一条独立数据记录,一个请求返回的结果页里最多包含
page_size个Document。
额外需求:修改/删除部分文档
如果需要修改或删除导出后的部分文档,可以用以下方法:
修改文档
# 示例:修改第一条文档的price字段 doc_to_update = all_docs[0] doc_to_update["price"] = 49.99 search_client.upload_documents(documents=[doc_to_update])
删除文档
# 示例:删除库存为0的文档(假设主键字段是"id") doc_ids_to_delete = [doc["id"] for doc in filtered_docs if doc["stock"] == 0] search_client.delete_documents(documents=[{"id": doc_id} for doc_id in doc_ids_to_delete])
内容的提问来源于stack exchange,提问作者newbie101
相关产品推荐
相关产品推荐

