You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

Elasticsearch 6.8.5是否支持忽略关键词中的空格进行分析?

How to Ignore Spaces in Keyword Fields for Queries in Elasticsearch 6.8.5

Hey there! Let's break down what's happening and how to fix it.

First, the root issue: keyword fields in Elasticsearch are not analyzed by default. That means they treat the entire string as a single, exact match. So when your data has "100 x 2c ABSD", the keyword field stores it exactly like that—so only queries for "100 x" will hit, while "100x" (no space) is treated as a completely different value.

To make both "100 x" and "100x" return your target documents, you'll need to create a custom analyzed field that strips out spaces, while still keeping the original keyword field if you need exact matches later. Here's how to do it, depending on whether you're setting up a new index or modifying an existing one.

Option 1: Set Up a New Index with a Custom Analyzer

If you're starting fresh, define an index with a custom analyzer that removes spaces, then map your field to use this analyzer for a sub-field.

  1. Create the index with the custom analysis settings:
PUT /your_index_name
{
  "settings": {
    "analysis": {
      "char_filter": {
        "strip_spaces": {
          "type": "pattern_replace",
          "pattern": "\\s+",
          "replacement": ""
        }
      },
      "analyzer": {
        "space_free_keyword": {
          "type": "custom",
          "char_filter": ["strip_spaces"],
          "tokenizer": "keyword"
        }
      }
    }
  },
  "mappings": {
    "_doc": {  // Note: ES 6.x uses document types like "_doc" (removed in 7.x+)
      "properties": {
        "product_code": {  // Replace with your actual field name
          "type": "text",
          "fields": {
            "no_space": {
              "type": "keyword",
              "analyzer": "space_free_keyword"
            },
            "raw": {
              "type": "keyword"  // Keep original keyword for exact matches
            }
          }
        }
      }
    }
  }
}

This analyzer uses a pattern_replace char filter to remove all spaces, then the keyword tokenizer ensures the entire modified string stays as a single token. So "100 x 2c ABSD" becomes "100x2cABSD" in the no_space sub-field.

  1. Query to match both spaced and unspaced terms:
    You can either query both the raw and no-space fields, or normalize your query term (remove spaces) and hit the no-space field. Here's a flexible bool query that covers both cases:
GET /your_index_name/_search
{
  "query": {
    "bool": {
      "should": [
        {"match": {"product_code.raw": "100 x"}},  // Match exact spaced term
        {"wildcard": {"product_code.no_space": "100x*"}}  // Match unspaced prefix
      ],
      "minimum_should_match": 1
    }
  }
}

Or if you pre-process your query to remove spaces, you can just use a prefix query on the no_space field for better performance:

GET /your_index_name/_search
{
  "query": {
    "prefix": {
      "product_code.no_space": "100x"
    }
  }
}

Option 2: Modify an Existing Index

If you already have data in your index, you'll need to update the index settings, add the new sub-field, then reprocess your documents.

  1. Update the index settings to add the custom analyzer:
PUT /your_existing_index/_settings
{
  "analysis": {
    "char_filter": {
      "strip_spaces": {
        "type": "pattern_replace",
        "pattern": "\\s+",
        "replacement": ""
      }
    },
    "analyzer": {
      "space_free_keyword": {
        "type": "custom",
        "char_filter": ["strip_spaces"],
        "tokenizer": "keyword"
      }
    }
  }
}
  1. Update the mapping to add the no_space sub-field:
PUT /your_existing_index/_mapping/_doc
{
  "properties": {
    "product_code": {
      "type": "text",
      "fields": {
        "no_space": {
          "type": "keyword",
          "analyzer": "space_free_keyword"
        },
        "raw": {
          "type": "keyword"
        }
      }
    }
  }
}
  1. Reprocess existing documents to populate the new sub-field:
    Run an _update_by_query to reindex all documents, which will generate the no_space values for existing data:
POST /your_existing_index/_update_by_query
{
  "query": {
    "match_all": {}
  }
}

You could technically use a script to strip spaces at query time, but this is not recommended for production—it's slow, especially with large datasets, because it processes every document on the fly. For completeness, here's what that looks like:

GET /your_existing_index/_search
{
  "query": {
    "bool": {
      "filter": {
        "script": {
          "script": "doc['product_code.keyword'].value.replaceAll('\\\\s+', '').startsWith('100x')"
        }
      }
    }
  }
}

Stick to the custom analyzer approach for better performance and scalability.


内容的提问来源于stack exchange,提问作者Xavi Nguyen

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.14 09:08:55