You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何在Solr中搜索含数字与特殊字符的词汇?splitOnNumerics设置后仍有异常

Fixing Search for Terms with Special Characters When splitOnNumerics is 0

Great question—this is a common gotcha with tokenization and special characters in search engines. Let’s break down why this is happening and how to fix it.

Root Cause

Setting splitOnNumerics: 0 tells the tokenizer not to split numbers from adjacent letters (hence why "90s" and "omega30" work), but it doesn’t change how non-word characters like " or ) are handled. Most standard tokenizers (like Elasticsearch’s standard tokenizer) split tokens at these non-word characters by default.

For example, when you index "80\"", the tokenizer will split it into just "80" (discarding the quote), so even if you escape the quote in your search query, there’s no matching token to find.

Solutions

Here are three actionable fixes depending on your use case:

1. Create a Custom Tokenizer to Include Special Characters

If you want terms like "80\"" or "40)" to be treated as single searchable tokens, define a custom analyzer that includes your desired special characters in the allowed token set.

For example, in Elasticsearch, you’d set up a custom pattern tokenizer that splits on anything not in your allowed character set (letters, numbers, ", and )):

{
  "settings": {
    "analysis": {
      "analyzer": {
        "mixed_special_analyzer": {
          "tokenizer": "mixed_special_tokenizer"
        }
      },
      "tokenizer": {
        "mixed_special_tokenizer": {
          "type": "pattern",
          "pattern": "[^a-zA-Z0-9\"()]+", // Split on non-allowed characters
          "group": 0
        }
      }
    }
  }
}

Apply this analyzer to your target field, and when you index "80\"" or "40)", they’ll be stored as full tokens that your escaped queries can match.

2. Use a Keyword Sub-Field for Exact Matches

If you don’t want to overhaul your main analyzer, add a keyword sub-field to your mapping. This field stores the exact, un-tokenized string, perfect for precise searches:

{
  "mappings": {
    "properties": {
      "your_search_field": {
        "type": "text",
        "analyzer": "standard",
        "split_on_numerics": 0,
        "fields": {
          "keyword": {
            "type": "keyword",
            "ignore_above": 256
          }
        }
      }
    }
  }
}

Then query the keyword field directly with your escaped term:

{
  "query": {
    "term": {
      "your_search_field.keyword": "80\""
    }
  }
}

Just make sure to escape characters correctly (e.g., in JSON, you’ll need to write "80\\\"" to pass the escaped quote).

3. Verify Tokenization and Adjust Queries

If you can’t modify your index, first check how your terms are being tokenized using your search engine’s analyze API. For Elasticsearch:

GET _analyze
{
  "analyzer": "standard",
  "split_on_numerics": 0,
  "text": "80\""
}

If the result is just the token "80", you can use a wildcard query to match partial terms (note: this is less efficient than token-based searches):

{
  "query": {
    "wildcard": {
      "your_search_field": "80*"
    }
  }
}

Or combine a term query for "80" with a filter for the presence of the quote character if your engine supports it.

Final Notes

Always double-check that your escaped characters are being passed correctly to the search engine—different contexts (like JSON strings) may require double-escaping (e.g., \\ instead of \).

内容的提问来源于stack exchange,提问作者metylbk

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.05.22 07:42:31