You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何使用Python将JSON文件导入Elasticsearch并按ID分索引存储

问题说明

你当前代码无法按ID拆分索引的核心原因是索引名被硬编码为固定值elasticsearch_index,同时代码还存在缩进错误、ES版本兼容问题、冗余序列化的小问题,调整后即可实现按单条数据的ID独立写入对应索引的需求。

调整后单条写入实现代码

前置说明:默认你的JSON数组中每条记录都带有唯一标识字段id,如果实际ID字段名不同,替换代码中对应的键名即可

import json
from elasticsearch import Elasticsearch

# 初始化ES连接,根据自身集群配置调整参数
es = Elasticsearch(
    [{'host': 'localhost', 'port': 9200}],
    # 开启账号认证的集群解开下方注释填写对应信息
    # http_auth=('es用户名', 'es密码'),
    # HTTPS协议集群解开下方注释跳过证书校验(生产环境不建议使用)
    # verify_certs=False
)

# 提前校验服务连通性
if not es.ping():
    raise ConnectionError("Elasticsearch连接失败,请检查服务状态和连接配置")

with open('js.json', 'r', encoding='utf-8') as raw_data:
    json_docs = json.load(raw_data)
    for json_doc in json_docs:
        # 提取数据ID并做格式处理,符合ES索引名规范
        raw_doc_id = str(json_doc['id'])
        # ES索引名要求全小写,不能包含空格、特殊符号,做基础替换
        format_doc_id = raw_doc_id.lower().replace(' ', '_').replace('/', '_').replace('\\', '_')
        target_index = f"business_data_{format_doc_id}"
        
        # 写入对应索引,ES7.0+版本无需指定doc_type参数
        es.index(
            index=target_index,
            id=format_doc_id,
            body=json_doc
        )
        print(f"数据已成功写入索引: {target_index}")
大数据量优化方案

如果你的JSON文件数据量超过1000条,单条循环写入性能较差,可以使用ES官方提供的批量写入接口,性能可以提升5~10倍,代码如下:

import json
from elasticsearch import Elasticsearch, helpers

es = Elasticsearch([{'host': 'localhost', 'port': 9200}])

def doc_generator():
    with open('js.json', 'r', encoding='utf-8') as raw_data:
        json_docs = json.load(raw_data)
        for json_doc in json_docs:
            raw_doc_id = str(json_doc['id'])
            format_doc_id = raw_doc_id.lower().replace(' ', '_').replace('/', '_').replace('\\', '_')
            yield {
                "_index": f"business_data_{format_doc_id}",
                "_id": format_doc_id,
                "_source": json_doc
            }

success_count, fail_count = helpers.bulk(es, doc_generator(), stats_only=True)
print(f"批量写入完成,成功写入{success_count}条,失败{fail_count}条")
注意事项
  • 索引名规范:Elasticsearch强制要求索引名为全小写,不能包含空格、/、\、*、?、"、<、>、|、#等特殊字符,代码中仅做了最常见的特殊字符替换,如果你的ID中包含其他违规字符,需要补充替换规则避免写入报错
  • 版本适配:如果你使用的是6.x及更低版本的Elasticsearch,需要在es.index或者批量生成的action中增加_type: 'raw_data'参数,7.0及以上版本已经废弃类型概念,不需要传该参数
  • 索引前缀:代码中索引用了business_data_作为固定前缀,你可以根据自身业务场景替换成更易识别的前缀,方便后续在Kibana中统一管理索引模式

内容的提问来源于stack exchange,提问作者Ashutosh Jasrotia

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.09.20 16:15:49