关于Elasticsearch多值字段position_increment_gap的技术问询
position_increment_gap in Elasticsearch Multivalue Fields Great question—let’s break down exactly how this parameter works, using your example to make it concrete.
What position_increment_gap does
When you have a multivalue field (like your names array), Elasticsearch treats each element in the array as a separate chunk of text. By default, it inserts a large "gap" (default value: 100) between the last term of one chunk and the first term of the next. This is a guardrail to prevent accidental phrase matches across unrelated elements in the array.
Your example, step by step
Let’s walk through how your document is indexed, depending on the gap setting:
Default gap (100) scenario
If you used Elasticsearch’s default mapping, here’s how terms are assigned positions:
"John Abraham"gets split intoJohn(position 0) andAbraham(position 1)- A gap of 100 is added between this chunk and the next. So the first term of the next chunk starts at
1 + 1 + 100 = 102 "Lincoln Smith"becomesLincoln(position 102) andSmith(position 103)- Another gap of 100 applies to the third chunk, starting at position 204
With this setup, your phrase query "Abraham Lincoln" would not match—since Abraham is at position 1 and Lincoln is at 102, their positions are 101 apart, which is way more than the "consecutive" requirement for a phrase match.
Gap set to 0 (your matching scenario)
Since you said the query matched as expected, you must have followed the official docs example and set position_increment_gap: 0 in your index mapping. Here’s how that changes things:
"John Abraham"still maps toJohn(0) andAbraham(1)- No gap is added, so the next chunk starts at
1 + 1 + 0 = 2 "Lincoln Smith"becomesLincoln(position 2) andSmith(position 3)
Now Abraham (position 1) is immediately followed by Lincoln (position 2)—exactly what the phrase query is looking for, hence the match.
Why this parameter matters
- Default gap (100): Use this when you don’t want phrases to span across array elements. For example, if your
namesarray had ["John Doe", "Jane Smith"], you wouldn’t want a query for "Doe Jane" to return a false positive. - Gap set to 0: Use this only when you intend for phrases to cross array elements. This is useful for cases like a list of keywords that form a continuous phrase, or when you want to treat the entire array as a single block of text.
Test it yourself
To see the position assignments firsthand, use the _analyze API with your field’s mapping:
POST /my_index/_analyze { "field": "names", "text": ["John Abraham", "Lincoln Smith"] }
Check the position values in the response—this will show you exactly how the gap is applied between the two text chunks.
内容的提问来源于stack exchange,提问作者user2434291

