如何从文本提取运动名称并在OpenSearch中精准匹配查询?
问题描述
我正在使用OpenSearch,现有一段包含多个运动名称的长文本,需要从中提取运动名称并在OpenSearch索引中匹配对应文档。输入文本格式多样,包含大小写字母、数字及特殊字符,运动名称无固定格式。
示例输入文本:
I will make a good 10 push-ups and Dumbbell Deficit Push-up
索引中的数据示例:
[ { "id": 2, "name": "Ankle Circles" }, { "id": 3, "name": "Barbell Deep Squat" }, { "id": 10, "name": "Push-ups" }, { "id": 11, "name": "Sit-up" }, { "id": 12, "name": "Air Squats" }, { "id": 13, "name": "Dumbbell Deficit Push-up" }, { "id": 14, "name": "Pretzel Stretch" }, { "id": 15, "name": "Cobra Stretch" }, { "id": 20, "name": "Push-ups with Elevated Feet" } ]
当前查询代码:
SearchResponse<ExerciseOSDto> searchResponse = openSearchClient.search( s -> s.index("exercises") .query(new Query.Builder().match( new MatchQuery.Builder() .field("name") .query(new FieldValue.Builder() .stringValue(payload.getText()).build()) .operator(Operator.Or) .build()) .build()), ExerciseOSDto.class);
当前查询会返回所有包含up/ups/push的运动(比如id=11、20的文档),但我只希望获取id为10和13的文档。请问从文本提取运动名称并在OpenSearch中搜索的最优方案是什么?
最优解决方案
核心思路是先从输入文本中提取精准的运动候选词,再通过OpenSearch的精准匹配查询获取目标文档,避免分词导致的模糊匹配问题。以下是具体实现步骤:
1. 预处理输入文本,提取候选运动名称
先对杂乱的输入文本做清洗和提取:
- 去除数字、无关特殊字符(比如示例中的
10),保留字母、空格和运动名称常用的连字符- - 按空格拆分文本后,尝试组合成完整的运动名称(比如从拆分后的单词中组合出
push-ups和Dumbbell Deficit Push-up) - 统一转为小写,消除大小写差异对匹配的影响
示例预处理后得到候选列表:["push-ups", "dumbbell deficit push-up"]
2. 优化OpenSearch索引字段映射(推荐)
给name字段添加keyword子字段,配合小写归一化器,实现大小写不敏感的精准匹配:
# 创建小写归一化器 PUT /_settings { "index": { "analysis": { "normalizer": { "lowercase_normalizer": { "type": "custom", "filter": ["lowercase"] } } } } } # 更新索引映射 PUT /exercises/_mapping { "properties": { "name": { "type": "text", "fields": { "keyword": { "type": "keyword", "ignore_above": 256, "normalizer": "lowercase_normalizer" } } } } }
这个配置既保留原text字段的分词能力,又新增了支持精准匹配的name.keyword字段。
3. 使用精准匹配查询获取目标文档
方案A:terms批量匹配候选词
直接用terms查询匹配预处理后的候选列表,确保只有完全匹配的文档被返回:
import java.util.Arrays; import java.util.stream.Collectors; import org.opensearch.client.opensearch.core.SearchResponse; import org.opensearch.client.opensearch._types.FieldValue; import org.opensearch.client.opensearch._types.query_dsl.Query; // 预处理得到的候选运动名称列表 List<String> exerciseCandidates = Arrays.asList("push-ups", "dumbbell deficit push-up"); SearchResponse<ExerciseOSDto> searchResponse = openSearchClient.search( s -> s.index("exercises") .query(q -> q.terms(t -> t .field("name.keyword") .terms(terms -> terms.value(exerciseCandidates.stream() .map(FieldValue::of) .collect(Collectors.toList()))))), ExerciseOSDto.class);
方案B:match_phrase短语匹配(适合格式轻微差异场景)
如果需要兼容少量格式差异(比如push-ups和push up),可以用match_phrase配合slop:0实现精准短语匹配:
import java.util.Arrays; import org.opensearch.client.opensearch.core.SearchResponse; import org.opensearch.client.opensearch._types.query_dsl.BoolQuery; import org.opensearch.client.opensearch._types.query_dsl.Query; // 预处理得到的候选运动名称列表 List<String> exerciseCandidates = Arrays.asList("push-ups", "dumbbell deficit push-up"); // 构建布尔查询,每个候选词对应一个match_phrase子查询 BoolQuery.Builder boolQuery = new BoolQuery.Builder(); for (String candidate : exerciseCandidates) { boolQuery.should(q -> q.matchPhrase(mp -> mp .field("name") .query(candidate) .slop(0))); } SearchResponse<ExerciseOSDto> searchResponse = openSearchClient.search( s -> s.index("exercises") .query(q -> q.bool(boolQuery.build())), ExerciseOSDto.class);
4. 提升候选词提取的准确性
如果运动名称集合是固定的,建议维护一个运动名称字典:
- 把索引中所有运动名称转成小写后存入字典
- 用滑动窗口扫描输入文本,直接匹配字典中的完整名称
- 这种方式能避免拆分组合带来的错误,提取精度更高
内容的提问来源于stack exchange,提问作者Taras Vovk
相关产品推荐
相关产品推荐

