You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Opensearch中支持模糊匹配的cross_fields替代方案

问题

我需要实现一个面向学生的职位搜索功能,要求基于职位关键词查询工作。目前用cross_fields类型的multi_match查询实现了理想的排序结果(多字段匹配权重累加,跨字段满足所有查询词匹配),但cross_fields不支持模糊匹配,无法处理拼写错误的场景。

测试数据集

POST _bulk
{"index":{"_index":"duarte-search-role","_id":"Job 1"}}
{"title":"Marine Biologist 1","overview":"Marine Biologist","opportunity_type_name":"Graduate Job","expired":false}
{"index":{"_index":"duarte-search-role","_id":"Job 2"}}
{"title":"Marine Biologist 2","overview":"No keyword","opportunity_type_name":"Graduate Job","expired":false}
{"index":{"_index":"duarte-search-role","_id":"Job 3"}}
{"title":"Comparison Job 3","overview":"Marine Biologist","opportunity_type_name":"Graduate Job","expired":false}
{"index":{"_index":"duarte-search-role","_id":"Job 4"}}
{"title":"Marine Biologist 4","overview":"Marine Biologist","opportunity_type_name":"Internship","expired":false}
{"index":{"_index":"duarte-search-role","_id":"Job 5"}}
{"title":"Marine Biologist 5","overview":"No keyword","opportunity_type_name":"Internship","expired":false}
{"index":{"_index":"duarte-search-role","_id":"Job 6"}}
{"title":"Comparison Job 6","overview":"Marine Biologist","opportunity_type_name":"Internship","expired":false}

预期搜索逻辑

  • 输入Marine Biologist:Job 1(标题+概述均匹配)排第一,Job 3(概述匹配)次之,Job 2(仅标题匹配)随后,以此类推
  • 输入Graduate Marine Biologist:Job 1(标题+概述匹配关键词,机会类型匹配Graduate)排第一,Job 2次之,以此类推
  • 输入Marine Biologist Internship:Job 4(标题+概述匹配关键词,机会类型匹配Internship)排第一,Job 5次之,以此类推

当前实现的查询(无模糊匹配)

GET /duarte-search-role/_search?search_type=dfs_query_then_fetch
{ 
  "query": {
    "bool": {
      "must": [
        {
          "multi_match": {
            "query": "Marine Biologist Internship",
            "fields": [
              "title^100",
              "overview^50",
              "opportunity_type_name^30"
            ],
            "operator": "and",
            "type": "cross_fields",
            "tie_breaker": 1
          }
        }
      ],
      "filter": [
        {
          "term": {
            "expired": false
          }
        }
      ]
    }
  },
  "sort": [
    {
      "_score": {
        "order": "desc"
      }
    },
    {
      "application_close_date": {
        "order": "asc"
      }
    }
  ],
  "from": 0,
  "size": 8
}

核心需求

保留原有排序逻辑(多字段权重累加、跨字段满足所有查询词匹配)的同时,支持模糊匹配以处理拼写错误。


解决方案

由于cross_fields类型不支持模糊匹配,可通过以下两种方案替代,同时保留原有排序逻辑:

方案1:使用query_string查询(无需拆分查询词)

query_string支持跨字段模糊匹配,且可通过default_operator: AND确保所有查询词都被匹配,得分会根据字段权重累加,符合原有排序逻辑。

GET /duarte-search-role/_search?search_type=dfs_query_then_fetch
{ 
  "query": {
    "bool": {
      "must": [
        {
          "query_string": {
            "query": "Marine~ Biologist~ Internship~",
            "fields": [
              "title^100",
              "overview^50",
              "opportunity_type_name^30"
            ],
            "default_operator": "AND",
            "fuzziness": "AUTO"
          }
        }
      ],
      "filter": [
        {
          "term": {
            "expired": false
          }
        }
      ]
    }
  },
  "sort": [
    {
      "_score": {
        "order": "desc"
      }
    },
    {
      "application_close_date": {
        "order": "asc"
      }
    }
  ],
  "from": 0,
  "size": 8
}

说明

  • 每个词后的~标记开启模糊匹配,也可通过全局fuzziness: "AUTO"统一设置模糊度(自动根据词长调整编辑距离)
  • default_operator: AND确保所有查询词都必须在某个字段匹配到,模拟cross_fields的operator: AND效果
  • 字段权重设置与原有查询一致,得分会根据匹配字段的权重累加,保留原有排序逻辑

方案2:拆分查询词,使用bool+multi_match组合(更灵活)

将用户输入的查询拆分为单个关键词,每个词对应一个multi_match模糊查询,放入bool的must子句中,确保所有词都被匹配,同时保留字段权重和累加得分逻辑。

GET /duarte-search-role/_search?search_type=dfs_query_then_fetch
{ 
  "query": {
    "bool": {
      "must": [
        {
          "multi_match": {
            "query": "Marine",
            "fields": ["title^100", "overview^50", "opportunity_type_name^30"],
            "fuzziness": "AUTO"
          }
        },
        {
          "multi_match": {
            "query": "Biologist",
            "fields": ["title^100", "overview^50", "opportunity_type_name^30"],
            "fuzziness": "AUTO"
          }
        },
        {
          "multi_match": {
            "query": "Internship",
            "fields": ["title^100", "overview^50", "opportunity_type_name^30"],
            "fuzziness": "AUTO"
          }
        }
      ],
      "filter": [
        {
          "term": {
            "expired": false
          }
        }
      ]
    }
  },
  "sort": [
    {
      "_score": {
        "order": "desc"
      }
    },
    {
      "application_close_date": {
        "order": "asc"
      }
    }
  ],
  "from": 0,
  "size": 8
}

说明

  • 需要在应用层将用户输入的查询拆分为单个词(可通过空格分割,注意处理连续空格等情况)
  • 每个multi_match针对单个词进行跨字段模糊匹配,must子句确保所有词都必须匹配
  • 字段权重与原有查询一致,得分由多个匹配字段的权重累加,完全保留原有排序逻辑
  • 相比query_string,此方案更灵活,可针对不同词设置不同的模糊度或字段范围

内容的提问来源于stack exchange,提问作者kyuubi

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.07.15 18:17:02