如何用Karate API动态识别重复值并关联字段与文档信息
问题内容
输入示例
[ { "fieldName": "fieldOne", "normalJson": "{\"nameRawText\":\"Man\",\"addressRawText\":\"23 Third \",\"cityRawText\":\"LONDON\",\"stateRawText\":null,\"countryRawText\":\"United Kingdom\",\"rawText\":\"Man 23 Third, LONDON United Kingdom\"}", "documentName": "documentOne" }, { "fieldName": "fieldTwo", "normalJson": "{\"nameRawText\":\"Man\",\"addressRawText\":\"23 Third SE55\",\"cityRawText\":\"LONDON\",\"stateRawText\":null,\"countryRawText\":\"United Kingdom\",\"rawText\":\"Man 23 Third, LONDON United Kingdom\"}", "documentName": "documentOne" }, { "fieldName": "fieldThree", "normalJson": "{\"nameRawText\":\"Mick\",\"addressRawText\":null,\"cityRawText\":null,\"stateRawText\":null,\"countryRawText\":\"India\",\"rawText\":\"Mick India\"}", "documentName": "documentOne" }, { "fieldName": "fieldSix", "normalJson": "{\"nameRawText\":\"Mick\",\"addressRawText\":null,\"cityRawText\":null,\"stateRawText\":null,\"countryRawText\":\"India\",\"rawText\":\"Mick India\"}", "documentName": "documentOne" }, { "fieldName": "fieldOne", "normalJson": "{\"nameRawText\":\"Man\",\"addressRawText\":\"23 Third \",\"cityRawText\":\"LONDON\",\"stateRawText\":null,\"countryRawText\":\"United Kingdom\",\"rawText\":\"Man 23 Third, LONDON United Kingdom\"}", "documentName": "documentTwo" }, { "fieldName": "fieldFour", "normalJson": "{\"nameRawText\":\"Man\",\"addressRawText\":\"23 Third \",\"cityRawText\":\"LONDON\",\"stateRawText\":null,\"countryRawText\":\"United Kingdom\",\"rawText\":\"Man 23 Third, LONDON United Kingdom\"}", "documentName": "documentTwo" } ]
预期输出
[ { "nameRawText": "Man", "fieldName": "fieldOne,fieldTwo,fieldFour", "documentName": "documentOne,documentTwo" }, { "nameRawText": "Mick", "fieldName": "fieldThree,fieldSix", "documentName": "documentOne" } ]
注:原预期输出存在JSON语法错误(重复键),上述为修正后的合法格式。
测试场景
识别normalJson中nameRawText、cityRawText或countryRawText的重复值,输出重复值及其对应键、fieldName和documentName。
需求说明
需提供动态处理上述场景的方案,满足以下要求:
- 载荷包含可变数量的文档与字段;
- 已尝试
karate.filter、karate.jsonPath、HashMap未得到预期结果; - IDE不支持
karate.distinct方法。
解决方案
核心思路
通过Karate原生循环+Java集合实现分组聚合,全程规避karate.distinct依赖,同时支持动态数量的输入条目。
步骤1:解析嵌套JSON字符串
先把每个条目中的normalJson字符串转为可操作的JSON对象:
* def parsedItems = karate.map(input, function(item) { item.normalJson = karate.fromJson(item.normalJson); return item; })
步骤2:按目标字段分组聚合
创建Map存储分组数据,用LinkedHashSet保证字段值去重且保留顺序:
* def groupMap = {} // 定义需要检测重复的目标字段 * def targetFields = ['nameRawText', 'cityRawText', 'countryRawText'] * karate.forEach(parsedItems, function(item) { karate.forEach(targetFields, function(field) { var fieldValue = item.normalJson[field]; if (fieldValue == null) return; // 跳过空值 // 用"字段名:字段值"作为分组唯一标识 var groupKey = field + ':' + fieldValue; if (!groupMap[groupKey]) { groupMap[groupKey] = { [field]: fieldValue, fieldName: new java.util.LinkedHashSet(), documentName: new java.util.LinkedHashSet() }; } // 加入当前条目的字段和文档信息 groupMap[groupKey].fieldName.add(item.fieldName); groupMap[groupKey].documentName.add(item.documentName); }) })
步骤3:转换为预期输出格式
将集合转为逗号分隔的字符串,最终整理为数组:
* def result = karate.map(Object.values(groupMap), function(entry) { entry.fieldName = entry.fieldName.join(','); entry.documentName = entry.documentName.join(','); return entry; })
关键说明
- 支持任意数量的输入文档和字段,无需硬编码;
LinkedHashSet同时实现去重和顺序保留,符合业务需求;- 完全规避
karate.distinct依赖,兼容受限IDE环境。
内容的提问来源于stack exchange,提问作者sri
相关产品推荐
相关产品推荐

