You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

关于Elasticsearch多值字段position_increment_gap的技术问询

Understanding position_increment_gap in Elasticsearch Multivalue Fields

Great question—let’s break down exactly how this parameter works, using your example to make it concrete.

What position_increment_gap does

When you have a multivalue field (like your names array), Elasticsearch treats each element in the array as a separate chunk of text. By default, it inserts a large "gap" (default value: 100) between the last term of one chunk and the first term of the next. This is a guardrail to prevent accidental phrase matches across unrelated elements in the array.

Your example, step by step

Let’s walk through how your document is indexed, depending on the gap setting:

Default gap (100) scenario

If you used Elasticsearch’s default mapping, here’s how terms are assigned positions:

  • "John Abraham" gets split into John (position 0) and Abraham (position 1)
  • A gap of 100 is added between this chunk and the next. So the first term of the next chunk starts at 1 + 1 + 100 = 102
  • "Lincoln Smith" becomes Lincoln (position 102) and Smith (position 103)
  • Another gap of 100 applies to the third chunk, starting at position 204

With this setup, your phrase query "Abraham Lincoln" would not match—since Abraham is at position 1 and Lincoln is at 102, their positions are 101 apart, which is way more than the "consecutive" requirement for a phrase match.

Gap set to 0 (your matching scenario)

Since you said the query matched as expected, you must have followed the official docs example and set position_increment_gap: 0 in your index mapping. Here’s how that changes things:

  • "John Abraham" still maps to John (0) and Abraham (1)
  • No gap is added, so the next chunk starts at 1 + 1 + 0 = 2
  • "Lincoln Smith" becomes Lincoln (position 2) and Smith (position 3)

Now Abraham (position 1) is immediately followed by Lincoln (position 2)—exactly what the phrase query is looking for, hence the match.

Why this parameter matters

  • Default gap (100): Use this when you don’t want phrases to span across array elements. For example, if your names array had ["John Doe", "Jane Smith"], you wouldn’t want a query for "Doe Jane" to return a false positive.
  • Gap set to 0: Use this only when you intend for phrases to cross array elements. This is useful for cases like a list of keywords that form a continuous phrase, or when you want to treat the entire array as a single block of text.

Test it yourself

To see the position assignments firsthand, use the _analyze API with your field’s mapping:

POST /my_index/_analyze
{
  "field": "names",
  "text": ["John Abraham", "Lincoln Smith"]
}

Check the position values in the response—this will show you exactly how the gap is applied between the two text chunks.

内容的提问来源于stack exchange,提问作者user2434291

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.21 04:36:13