You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置AWS Data Prepper管道解决S3含数组JSON导入OpenSearch报错?

问题:AWS Data Prepper无法将含数组的JSON从S3导入OpenSearch

错误日志

2024-03-12T00:21:29.627 [Thread-11] ERROR org.opensearch.dataprepper.plugins.source.s3.S3ObjectWorker - Error reading from S3 object: s3ObjectReference=[bucketName=test2-8964, key=5a34c11a21ab22621b026ed00936aa425a2d75345585ddc06b380afd87ae7563.json]. Cannot construct instance of java.util.LinkedHashMap (although at least one Creator exists): no String-argument constructor/factory method to deserialize from String value ('dummystring') at [Source: (org.opensearch.dataprepper.plugins.source.s3.S3InputStream); line: 3, column: 9]

当前管道配置

version: "2"
test-pipeline:
  source:
    s3:
      codec:
        json:
      compression: "none"
      aws:
        region: "us-west-1"
        sts_role_arn: "arn:aws:iam::896497961826:role/test-pipeline"
      acknowledgments: true
      scan:
        scheduling:
          interval: PT30S
        buckets:
          - bucket:
              name: test2-8964
      delete_s3_objects_on_read: false
  processor:
    - date:
        destination: "@timestamp"
        from_time_received: true
  sink:
    - opensearch:
        hosts: ["https://search-test-6zryuwp5gchu72mzgpscbjx4qu.us-west-1.es.amazonaws.com"]
        index: "test-pipeline"
        aws:
          sts_role_arn: "arn:aws:iam::896497961826:role/test-pipeline"
          region: "us-west-1"
        dlq:
          s3:
            bucket: "test2-8964"
            region: "us-west-1"
            sts_role_arn: "arn:aws:iam::896497961826:role/test-pipeline"

S3中的JSON示例

{
    "array": [
        "dummystring"
    ]
}

解决方案

错误原因是Data Prepper错误地将JSON对象内数组的元素当作独立文档解析,而非将整个顶级JSON对象作为单个文档处理。修改S3源的JSON codec配置,明确指定解析规则即可解决:

  1. 更新管道配置中的S3 source部分,为json codec添加明确的mode和document_list_type参数:
source:
  s3:
    codec:
      json:
        mode: json
        document_list_type: none
    # 其余配置保持不变
  • mode: json:指定按单个JSON对象解析文件内容
  • document_list_type: none:告知Data Prepper不要将任何数组拆分为多个文档,将整个文件内容作为一个完整文档处理
  1. 移除之前添加的parse_json处理器,因为错误发生在源读取阶段,该处理器无法解决问题。

修改后的完整管道配置如下:

version: "2"
test-pipeline:
  source:
    s3:
      codec:
        json:
          mode: json
          document_list_type: none
      compression: "none"
      aws:
        region: "us-west-1"
        sts_role_arn: "arn:aws:iam::896497961826:role/test-pipeline"
      acknowledgments: true
      scan:
        scheduling:
          interval: PT30S
        buckets:
          - bucket:
              name: test2-8964
      delete_s3_objects_on_read: false
  processor:
    - date:
        destination: "@timestamp"
        from_time_received: true
  sink:
    - opensearch:
        hosts: ["https://search-test-6zryuwp5gchu72mzgpscbjx4qu.us-west-1.es.amazonaws.com"]
        index: "test-pipeline"
        aws:
          sts_role_arn: "arn:aws:iam::896497961826:role/test-pipeline"
          region: "us-west-1"
        dlq:
          s3:
            bucket: "test2-8964"
            region: "us-west-1"
            sts_role_arn: "arn:aws:iam::896497961826:role/test-pipeline"

内容的提问来源于stack exchange,提问作者Daring Cλlf

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 02:44:55