You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Python中结合Bulk API应用Elasticsearch新映射?

解决Elasticsearch Bulk导入时日期格式解析失败的问题

嘿,咱们先来拆解下你遇到的问题:从错误日志里能清楚看到,核心是datum字段的日期解析失败——你的源数据日期格式是2018-02-07 01:00:51,但Elasticsearch没办法识别它。

raise BulkIndexError('%i document(s) failed to index.' % len(errors), errors) BulkIndexError: (u'2 document(s) failed to index.', [{u'index': {u'status': 400, u'_type': u'kliks', u'_index': u'logstash-2018.02.07', u'error': {u'caused_by': {u'reason': u'Invalid format: "2018-02-07 01:00:51" is malformed at " 01:00:51"', u'type': u'illegal_argument_exception'}, u'reason': u'failed to parse [datum]', u'type': u'mapper_parsing_exception'}, ....

这是因为你在目标索引的映射里只把datum设为date类型,但没指定对应的日期格式。Elasticsearch默认用的是严格的ISO 8601格式(比如2018-02-07T01:00:51.000Z),而你的源数据是用空格分隔日期和时间的格式,自然就解析失败了。

解决办法:给日期字段指定匹配的格式

你只需要修改create_mapping函数里的datum字段定义,加上format参数,明确告诉Elasticsearch要识别的日期格式就行:

def create_mapping(es, idx, document_type):
    mymapping = {
        "mappings": {
            document_type: {
                "properties": {
                    "prijs": {"type": "integer"},
                    "datum": {
                        "type": "date",
                        "format": "yyyy-MM-dd HH:mm:ss"  # 完全匹配源数据的日期格式
                    },
                    "kilometerstand": {"type": "integer"}
                }
            }
        }
    }
    if es.indices.exists(index=idx):
        es.indices.delete(index=idx)
    es.indices.create(index=idx, body=mymapping)

要是源数据里可能存在多种日期格式,也可以用竖线分隔指定多个格式,让Elasticsearch依次尝试解析:

"datum": {
    "type": "date",
    "format": "yyyy-MM-dd HH:mm:ss||yyyy-MM-dd||ISO8601"
}

后续验证步骤

  1. 先删掉之前创建的目标索引(如果已经存在的话)
  2. 运行修改后的create_mapping函数重新创建索引
  3. 再执行copy_index批量导入,这次应该就能正常解析日期字段,不会再报BulkIndexError了

内容的提问来源于stack exchange,提问作者YNR

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.29 08:12:54