You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

使用AWS Java V2 SDK(Groovy)创建Kendra Web Crawler V2数据源报错排查

解决AWS Kendra Web Crawler V2数据源创建时的500错误

核心问题定位

你遇到的500错误大概率是模板JSON的字段类型不匹配导致的。Web Crawler V2的模板配置对数值类型字段有严格要求,你的JSON中多个数值字段被写成了字符串格式,这会触发后端解析失败。

具体修复步骤

1. 修正模板JSON的字段类型

将以下字段的字符串值改为数字(去掉引号):

  • rateLimit(原"300" → 300)
  • crawlDepth(原"5" → 5)
  • maxFileSize(原"50" → 50)
  • maxLinksPerUrl(原"1000" → 1000)

修正后的JSON示例:

{
  "connectionConfiguration": {
    "repositoryEndpointMetadata": {
      "authentication": "NoAuthentication",
      "siteMapUrls": [
        "https://www.my-domain.com/sitemap.xml"
      ]
    }
  },
  "repositoryConfigurations": {
    "webPage": {
      "fieldMappings": [
        {
          "dataSourceFieldName": "sourceUrl",
          "indexFieldName": "_source_uri",
          "indexFieldType": "STRING"
        },
        {
          "dataSourceFieldName": "category",
          "indexFieldName": "_category",
          "indexFieldType": "STRING"
        }
      ]
    },
    "attachment": {
      "fieldMappings": [
        {
          "dataSourceFieldName": "sourceUrl",
          "indexFieldName": "_source_uri",
          "indexFieldType": "STRING"
        },
        {
          "dataSourceFieldName": "category",
          "indexFieldName": "_category",
          "indexFieldType": "STRING"
        }
      ]
    }
  },
  "syncMode": "FULL_CRAWL",
  "additionalProperties": {
    "exclusionFileIndexPatterns": [],
    "exclusionURLCrawlPatterns": [],
    "exclusionURLIndexPatterns": [],
    "inclusionFileIndexPatterns": [],
    "inclusionURLCrawlPatterns": [
      ".*/gb-en/.*",
      ".*/de-en/.*"
    ],
    "inclusionURLIndexPatterns": [],
    "proxy": {},
    "rateLimit": 300,
    "crawlAllDomain": false,
    "crawlAttachments": true,
    "crawlDepth": 5,
    "maxFileSize": 50,
    "crawlSubDomain": true,
    "maxLinksPerUrl": 1000,
    "honorRobots": false
  },
  "type": "WEBCRAWLERV2"
}

2. 验证Document对象的构建

确保JSON文件读取和Document转换过程没有编码问题,建议添加异常捕获来确认JSON是否有效:

try {
    String documentJsonStr = Files.readString(Paths.get(getClass().getClassLoader().getResource("json/template-kendra-web-crawler.json").toURI()))
    // 验证JSON格式合法性
    new groovy.json.JsonSlurper().parseText(documentJsonStr)
    Document document = Document.fromString(documentJsonStr)
} catch (Exception e) {
    println "JSON解析失败: ${e.getMessage()}"
    throw e
}

3. 启用SDK调试日志排查

如果问题仍存在,启用AWS SDK的详细日志,查看完整的请求和响应内容,这能帮你定位更隐蔽的错误:
在你的日志配置中添加:

log4j.logger.software.amazon.awssdk=DEBUG
log4j.logger.com.amazonaws=DEBUG

4. 确认角色权限

确保Kendra角色拥有以下必要权限:

  • kendra:CreateDataSource
  • kendra:DescribeDataSource
  • 若使用VPC:ec2:DescribeSubnets、ec2:DescribeSecurityGroups
  • 访问目标网站的网络权限(如果在VPC内,需确保NAT网关配置正确)

修复后的完整代码示例

import software.amazon.awssdk.core.document.Document
import software.amazon.awssdk.services.kendra.KendraClient
import software.amazon.awssdk.services.kendra.model.DataSourceType
import software.amazon.awssdk.services.kendra.model.Tag
import groovy.json.JsonSlurper

import java.nio.file.Files
import java.nio.file.Paths

static void main(String[] args) {
    KendraClient kendraClient = getKendraClient()
    String indexId = getIndexId(kendraClient)
    String roleArn = getRoleArn(kendraClient)
    List<Tag> tags = getTags()
    
    String documentJsonStr
    try {
        documentJsonStr = Files.readString(Paths.get(getClass().getClassLoader().getResource("json/template-kendra-web-crawler.json").toURI()))
        // 提前验证JSON格式
        new JsonSlurper().parseText(documentJsonStr)
    } catch (Exception e) {
        println "加载或解析JSON模板失败: ${e.getMessage()}"
        return
    }
    Document document = Document.fromString(documentJsonStr)

    try {
        kendraClient.createDataSource { ds -> ds
            .indexId(indexId)
            .type(DataSourceType.TEMPLATE)
            .name("test-wc-v2-datasource")
            .description("Testing the web crawler v2 datasource")
            .roleArn(roleArn)
            .languageCode("en")
            .tags(tags)
            .schedule("cron(0 18 ? * MON-FRI *)")
            .vpcConfiguration { vpc -> vpc
                .subnetIds("subnet-12345", "subnet-67890")
                .securityGroupIds("sg-12345")
            }
            .configuration { config -> config
                .templateConfiguration { tmplConfig -> tmplConfig
                            .template(document)
                }
            }
        }
        println "数据源创建成功"
    } catch (Exception e) {
        println "创建失败: ${e.getMessage()}"
        // 打印完整堆栈跟踪以便排查
        e.printStackTrace()
    }
}

内容的提问来源于stack exchange,提问作者Daniel Mills

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.08 14:45:13