You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何用Python从Azure Search Index导出全部/部分文档至DataFrame/JSON?

从Azure Cognitive Search索引导出数据到DataFrame/JSON的Python方案

核心实现脚本

首先确保安装Azure Search的Python SDK:

pip install azure-search-documents azure-identity

以下是可直接复用的代码,支持导出全部文档或基于过滤条件导出,最终转为DataFrame或JSON:

from azure.search.documents import SearchClient
from azure.identity import DefaultAzureCredential
import pandas as pd

# 初始化SearchClient
service_endpoint = "你的搜索服务端点"
index_name = "你的索引名称"
# 两种认证方式选其一:用默认身份认证,或直接传入API密钥
credential = DefaultAzureCredential()
# credential = ApiKeyCredential(api_key="你的API密钥")
search_client = SearchClient(endpoint=service_endpoint, index_name=index_name, credential=credential)

def export_search_documents(filter_query=None):
    """
    导出搜索索引文档
    :param filter_query: 过滤条件(遵循OData语法),示例:"category eq 'books'",不传则导出全部
    :return: 包含所有目标文档的列表
    """
    all_documents = []
    skip = 0
    page_size = 1000  # Azure Search单次请求最大返回1000条,为服务端硬限制
    while True:
        # search="*"匹配所有文档,select="*"返回所有字段
        results = search_client.search(
            search="*",
            filter=filter_query,
            select="*",
            top=page_size,
            skip=skip
        )
        page_docs = list(results)
        if not page_docs:
            break
        all_documents.extend(page_docs)
        skip += page_size
    return all_documents

# 示例1:导出全部文档并转为DataFrame
all_docs = export_search_documents()
df = pd.DataFrame(all_docs)
print(df.head())

# 示例2:按过滤条件导出并保存为JSON
filtered_docs = export_search_documents(filter_query="price lt 50 and category eq 'electronics'")
import json
with open("filtered_docs.json", "w", encoding="utf-8") as f:
    json.dump(filtered_docs, f, ensure_ascii=False, indent=2)

你关心的疑问解答

  1. page_size的1000限制是硬限制吗?
    是的,Azure Cognitive Search的Search API单次请求最多返回1000条结果,这是服务端设定的上限,无法突破。所以当索引中文档数量超过1000时,必须通过skip参数分页循环获取。

  2. page_size和Document的区别?

  • page_size(对应代码中的top参数)是单次API请求返回的文档数量上限;
  • Document就是你所说的索引中的"行项",每个Document是索引中的一条独立数据记录,一个请求返回的结果页里最多包含page_size个Document。

额外需求:修改/删除部分文档

如果需要修改或删除导出后的部分文档,可以用以下方法:

修改文档

# 示例:修改第一条文档的price字段
doc_to_update = all_docs[0]
doc_to_update["price"] = 49.99
search_client.upload_documents(documents=[doc_to_update])

删除文档

# 示例:删除库存为0的文档(假设主键字段是"id")
doc_ids_to_delete = [doc["id"] for doc in filtered_docs if doc["stock"] == 0]
search_client.delete_documents(documents=[{"id": doc_id} for doc_id in doc_ids_to_delete])

内容的提问来源于stack exchange,提问作者newbie101

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.14 14:10:13