You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

文档数组属性的过滤、聚合与分页实现方案咨询

实现灵活前端搜索器的Elasticsearch优化方案

问题背景

我正尝试为前端应用实现一个灵活的搜索器,相关细节如下:

文档结构

单条文档结构示例:

{
  "student": {
    "id": 1,
    "name": "Joe"
  },
  "city": {
    "id": 102,
    "name": "London"
  },
  "tags": [
    { "id": 33, "name": "football" },
    { "id": 34, "name": "basketball" },
    { "id": 35, "name": "music" },
    ...
  ],
  "skills": [
    { "id": 302, "name": "Active listening" },
    { "id": 23, "name": "Collaboration" },
    { "id": 34, "name": "Communication" },
    ...
  ]
}

已上传多条同结构文档,不同学生拥有不同标签,同一标签可分配给多个用户,技能的关联逻辑同理。

需求目标

前端有三个<Selects />组件(城市选择、标签选择、技能选择),当选择London时,需要:

  • 获取所有居住在伦敦的学生的标签列表
  • 支持通过<Select />输入文本过滤标签名称
  • 支持列表分页

当前配置

索引Mapping

{
  "mappings": {
    "properties": {
      "student": {
        "type": "nested",
        "properties": {
          "id": { "type": "integer" },
          "name": { "type": "keyword" }
        }
      },
      "city": {
        "type": "nested",
        "properties": {
          "id": { "type": "integer" },
          "name": { "type": "keyword" }
        }
      },
      "tags": {
        "type": "nested",
        "properties": { 
           "id": { "type": "integer" },
           "name": { 
             "type": "keyword",
             "normalizer": "lowercase_normalizer"
           }
        }
      },
      "skills": {
        "type": "nested",
        "properties": {
            "id": { "type": "integer" },
            "name": { 
              "type": "keyword",
              "normalizer": "lowercase_normalizer"
            }
        }
      }
    }
  }
}

分词器配置

为支持大小写不敏感的文本过滤,添加了normalizer:

"settings": {
  "analysis": {
    "normalizer": {
      "lowercase_normalizer": {
        "type": "custom",
        "char_filter": [],
        "filter": ["lowercase", "asciifolding"]
      }
    }
  }
}

尝试过的方案

  1. 支持分页但无法过滤标签名称的查询:
{
  "aggs": {
    "aggregatorField": {
      "nested": {
        "path": "tags"
      },
      "aggs": {
        "aggregator": {
          "composite": {
            "size": 11,
            "sources": [
              {
                "aggregator": {
                  "terms": {
                    "field": "tags.id"
                  }
                }
              }
            ]
          },
          "aggs": {
            "item": {
              "top_hits": {
                "size": 1,
                "sort": [
                  {
                    "tags.id": {
                      "order": "desc"
                    }
                  }
                ],
                "_source": {
                  "include": [
                    "tags"
                  ]
                }
              }
            }
          }
        }
      }
    }
  },
  "size": 0
}
  1. 后续实现了文本过滤,但查询不支持分页且逻辑冗余,考虑将标签、技能拆分到单独索引,实现类似SQL Join的关联操作,寻求更优方案。

优化方案

方案一:基于现有索引优化查询(推荐)

通过嵌套过滤+复合聚合组合,在不拆分索引的前提下实现所有需求:

最终查询示例

{
  "size": 0,
  "query": {
    "nested": {
      "path": "city",
      "query": {
        "term": {
          "city.name": "London"
        }
      }
    }
  },
  "aggs": {
    "filtered_tags": {
      "nested": {
        "path": "tags"
      },
      "aggs": {
        "tag_name_filter": {
          "filter": {
            "wildcard": {
              "tags.name": "*foot*" // 替换为Select组件输入的过滤文本,自动转为小写匹配normalizer
            }
          },
          "aggs": {
            "paginated_tags": {
              "composite": {
                "size": 10, // 每页展示数量
                "sources": [
                  {
                    "tag_id": {
                      "terms": {
                        "field": "tags.id"
                      }
                    }
                  }
                ],
                "after": { "tag_id": 33 } // 分页游标,传入上一页最后一个标签的ID实现翻页
              },
              "aggs": {
                "tag_info": {
                  "top_hits": {
                    "size": 1,
                    "_source": {
                      "include": ["tags.id", "tags.name"]
                    }
                  }
                }
              }
            }
          }
        }
      }
    }
  }
}

关键说明

  • 外层query先过滤出所有居住在London的文档
  • 嵌套聚合内通过filter实现标签名称的模糊匹配(可根据需求替换为match_phrase_prefix实现前缀匹配)
  • composite聚合天然支持分页,通过after参数传递上一页的聚合键值即可实现翻页
  • 利用lowercase_normalizer确保过滤时大小写不敏感

方案二:拆分索引实现关联查询

如果数据量极大或标签/技能更新频繁,可拆分三个独立索引:

  1. students:存储学生与城市关联信息
{
  "student_id": 1,
  "student_name": "Joe",
  "city_id": 102,
  "city_name": "London"
}
  1. tags:存储标签元数据
{
  "tag_id": 33,
  "tag_name": "football"
}
  1. student_tags:存储学生与标签的关联关系
{
  "student_id": 1,
  "tag_id": 33
}

查询流程

  1. 从students索引查询所有London学生的ID列表
  2. 用学生ID列表在student_tags索引中查询关联的标签ID集合
  3. 结合标签ID集合和过滤文本,在tags索引中查询目标标签并分页

优缺点

  • ✅ 优点:标签元数据无需重复存储,更新时只需修改tags索引,数据一致性更好;聚合查询逻辑更轻量
  • ❌ 缺点:需要多轮查询,增加前后端逻辑复杂度;无法利用Elasticsearch嵌套聚合的原生性能优化

总结

若当前数据量不大,优先选择方案一,通过优化现有查询即可满足需求,逻辑简洁且性能足够;若数据量达百万级以上且标签/技能更新频繁,再考虑方案二的索引拆分方案。

内容的提问来源于stack exchange,提问作者SnoberS

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.06.28 15:57:05