使用AWS Java V2 SDK(Groovy)创建Kendra Web Crawler V2数据源报错排查
解决AWS Kendra Web Crawler V2数据源创建时的500错误
核心问题定位
你遇到的500错误大概率是模板JSON的字段类型不匹配导致的。Web Crawler V2的模板配置对数值类型字段有严格要求,你的JSON中多个数值字段被写成了字符串格式,这会触发后端解析失败。
具体修复步骤
1. 修正模板JSON的字段类型
将以下字段的字符串值改为数字(去掉引号):
rateLimit(原"300" → 300)crawlDepth(原"5" → 5)maxFileSize(原"50" → 50)maxLinksPerUrl(原"1000" → 1000)
修正后的JSON示例:
{ "connectionConfiguration": { "repositoryEndpointMetadata": { "authentication": "NoAuthentication", "siteMapUrls": [ "https://www.my-domain.com/sitemap.xml" ] } }, "repositoryConfigurations": { "webPage": { "fieldMappings": [ { "dataSourceFieldName": "sourceUrl", "indexFieldName": "_source_uri", "indexFieldType": "STRING" }, { "dataSourceFieldName": "category", "indexFieldName": "_category", "indexFieldType": "STRING" } ] }, "attachment": { "fieldMappings": [ { "dataSourceFieldName": "sourceUrl", "indexFieldName": "_source_uri", "indexFieldType": "STRING" }, { "dataSourceFieldName": "category", "indexFieldName": "_category", "indexFieldType": "STRING" } ] } }, "syncMode": "FULL_CRAWL", "additionalProperties": { "exclusionFileIndexPatterns": [], "exclusionURLCrawlPatterns": [], "exclusionURLIndexPatterns": [], "inclusionFileIndexPatterns": [], "inclusionURLCrawlPatterns": [ ".*/gb-en/.*", ".*/de-en/.*" ], "inclusionURLIndexPatterns": [], "proxy": {}, "rateLimit": 300, "crawlAllDomain": false, "crawlAttachments": true, "crawlDepth": 5, "maxFileSize": 50, "crawlSubDomain": true, "maxLinksPerUrl": 1000, "honorRobots": false }, "type": "WEBCRAWLERV2" }
2. 验证Document对象的构建
确保JSON文件读取和Document转换过程没有编码问题,建议添加异常捕获来确认JSON是否有效:
try { String documentJsonStr = Files.readString(Paths.get(getClass().getClassLoader().getResource("json/template-kendra-web-crawler.json").toURI())) // 验证JSON格式合法性 new groovy.json.JsonSlurper().parseText(documentJsonStr) Document document = Document.fromString(documentJsonStr) } catch (Exception e) { println "JSON解析失败: ${e.getMessage()}" throw e }
3. 启用SDK调试日志排查
如果问题仍存在,启用AWS SDK的详细日志,查看完整的请求和响应内容,这能帮你定位更隐蔽的错误:
在你的日志配置中添加:
log4j.logger.software.amazon.awssdk=DEBUG log4j.logger.com.amazonaws=DEBUG
4. 确认角色权限
确保Kendra角色拥有以下必要权限:
kendra:CreateDataSourcekendra:DescribeDataSource- 若使用VPC:
ec2:DescribeSubnets、ec2:DescribeSecurityGroups - 访问目标网站的网络权限(如果在VPC内,需确保NAT网关配置正确)
修复后的完整代码示例
import software.amazon.awssdk.core.document.Document import software.amazon.awssdk.services.kendra.KendraClient import software.amazon.awssdk.services.kendra.model.DataSourceType import software.amazon.awssdk.services.kendra.model.Tag import groovy.json.JsonSlurper import java.nio.file.Files import java.nio.file.Paths static void main(String[] args) { KendraClient kendraClient = getKendraClient() String indexId = getIndexId(kendraClient) String roleArn = getRoleArn(kendraClient) List<Tag> tags = getTags() String documentJsonStr try { documentJsonStr = Files.readString(Paths.get(getClass().getClassLoader().getResource("json/template-kendra-web-crawler.json").toURI())) // 提前验证JSON格式 new JsonSlurper().parseText(documentJsonStr) } catch (Exception e) { println "加载或解析JSON模板失败: ${e.getMessage()}" return } Document document = Document.fromString(documentJsonStr) try { kendraClient.createDataSource { ds -> ds .indexId(indexId) .type(DataSourceType.TEMPLATE) .name("test-wc-v2-datasource") .description("Testing the web crawler v2 datasource") .roleArn(roleArn) .languageCode("en") .tags(tags) .schedule("cron(0 18 ? * MON-FRI *)") .vpcConfiguration { vpc -> vpc .subnetIds("subnet-12345", "subnet-67890") .securityGroupIds("sg-12345") } .configuration { config -> config .templateConfiguration { tmplConfig -> tmplConfig .template(document) } } } println "数据源创建成功" } catch (Exception e) { println "创建失败: ${e.getMessage()}" // 打印完整堆栈跟踪以便排查 e.printStackTrace() } }
内容的提问来源于stack exchange,提问作者Daniel Mills
相关产品推荐
相关产品推荐

