You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何从文本提取运动名称并在OpenSearch中精准匹配查询?

问题描述

我正在使用OpenSearch,现有一段包含多个运动名称的长文本,需要从中提取运动名称并在OpenSearch索引中匹配对应文档。输入文本格式多样,包含大小写字母、数字及特殊字符,运动名称无固定格式。

示例输入文本:

I will make a good 10 push-ups and Dumbbell Deficit Push-up

索引中的数据示例:

[
    {
        "id": 2,
        "name": "Ankle Circles"
    },
    {
        "id": 3,
        "name": "Barbell Deep Squat"
    },
    {
        "id": 10,
        "name": "Push-ups"
    },
    {
        "id": 11,
        "name": "Sit-up"
    },
    {
        "id": 12,
        "name": "Air Squats"
    },
    {
        "id": 13,
        "name": "Dumbbell Deficit Push-up"
    },
    {
        "id": 14,
        "name": "Pretzel Stretch"
    },
    {
        "id": 15,
        "name": "Cobra Stretch"
    },
    {
        "id": 20,
        "name": "Push-ups with Elevated Feet"
    }
]

当前查询代码:

SearchResponse<ExerciseOSDto> searchResponse = openSearchClient.search(
        s -> s.index("exercises")
            .query(new Query.Builder().match(
                    new MatchQuery.Builder()
                        .field("name")
                        .query(new FieldValue.Builder()
                            .stringValue(payload.getText()).build())
                        .operator(Operator.Or) 
                        .build())
                .build()), ExerciseOSDto.class);

当前查询会返回所有包含up/ups/push的运动(比如id=11、20的文档),但我只希望获取id为10和13的文档。请问从文本提取运动名称并在OpenSearch中搜索的最优方案是什么?

最优解决方案

核心思路是先从输入文本中提取精准的运动候选词,再通过OpenSearch的精准匹配查询获取目标文档,避免分词导致的模糊匹配问题。以下是具体实现步骤:

1. 预处理输入文本,提取候选运动名称

先对杂乱的输入文本做清洗和提取:

  • 去除数字、无关特殊字符(比如示例中的10),保留字母、空格和运动名称常用的连字符-
  • 按空格拆分文本后,尝试组合成完整的运动名称(比如从拆分后的单词中组合出push-ups和Dumbbell Deficit Push-up)
  • 统一转为小写,消除大小写差异对匹配的影响

示例预处理后得到候选列表:["push-ups", "dumbbell deficit push-up"]

2. 优化OpenSearch索引字段映射(推荐)

给name字段添加keyword子字段,配合小写归一化器,实现大小写不敏感的精准匹配:

# 创建小写归一化器
PUT /_settings
{
  "index": {
    "analysis": {
      "normalizer": {
        "lowercase_normalizer": {
          "type": "custom",
          "filter": ["lowercase"]
        }
      }
    }
  }
}

# 更新索引映射
PUT /exercises/_mapping
{
  "properties": {
    "name": {
      "type": "text",
      "fields": {
        "keyword": {
          "type": "keyword",
          "ignore_above": 256,
          "normalizer": "lowercase_normalizer"
        }
      }
    }
  }
}

这个配置既保留原text字段的分词能力,又新增了支持精准匹配的name.keyword字段。

3. 使用精准匹配查询获取目标文档

方案A:terms批量匹配候选词

直接用terms查询匹配预处理后的候选列表,确保只有完全匹配的文档被返回:

import java.util.Arrays;
import java.util.stream.Collectors;
import org.opensearch.client.opensearch.core.SearchResponse;
import org.opensearch.client.opensearch._types.FieldValue;
import org.opensearch.client.opensearch._types.query_dsl.Query;

// 预处理得到的候选运动名称列表
List<String> exerciseCandidates = Arrays.asList("push-ups", "dumbbell deficit push-up");

SearchResponse<ExerciseOSDto> searchResponse = openSearchClient.search(
    s -> s.index("exercises")
          .query(q -> q.terms(t -> t
              .field("name.keyword")
              .terms(terms -> terms.value(exerciseCandidates.stream()
                  .map(FieldValue::of)
                  .collect(Collectors.toList()))))),
    ExerciseOSDto.class);

方案B:match_phrase短语匹配(适合格式轻微差异场景)

如果需要兼容少量格式差异(比如push-ups和push up),可以用match_phrase配合slop:0实现精准短语匹配:

import java.util.Arrays;
import org.opensearch.client.opensearch.core.SearchResponse;
import org.opensearch.client.opensearch._types.query_dsl.BoolQuery;
import org.opensearch.client.opensearch._types.query_dsl.Query;

// 预处理得到的候选运动名称列表
List<String> exerciseCandidates = Arrays.asList("push-ups", "dumbbell deficit push-up");

// 构建布尔查询,每个候选词对应一个match_phrase子查询
BoolQuery.Builder boolQuery = new BoolQuery.Builder();
for (String candidate : exerciseCandidates) {
    boolQuery.should(q -> q.matchPhrase(mp -> mp
        .field("name")
        .query(candidate)
        .slop(0)));
}

SearchResponse<ExerciseOSDto> searchResponse = openSearchClient.search(
    s -> s.index("exercises")
          .query(q -> q.bool(boolQuery.build())),
    ExerciseOSDto.class);

4. 提升候选词提取的准确性

如果运动名称集合是固定的,建议维护一个运动名称字典:

  • 把索引中所有运动名称转成小写后存入字典
  • 用滑动窗口扫描输入文本,直接匹配字典中的完整名称
  • 这种方式能避免拆分组合带来的错误,提取精度更高

内容的提问来源于stack exchange,提问作者Taras Vovk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.26 02:57:41