You need to enable JavaScript to run this app.
优惠活动
大模型
产品
解决方案
定价
更多

如何配置ElasticSearch仅对URL字段前100字符生成N-gram词元

Solution: Truncate First, Then N-Gram

Absolutely you can optimize this! The key is to add a truncate character filter before your n-gram token filter. This will chop the URL down to its first 100 characters before generating any n-grams, eliminating all redundant tokens from the long tail of your URLs.

Here's a step-by-step configuration example:

1. Create an Index with Custom Analyzer

First, define an index that uses a custom analyzer combining truncation and n-gram processing:

PUT /url_index
{
  "settings": {
    "analysis": {
      "char_filter": {
        "truncate_url": {
          "type": "truncate",
          "length": 100
        }
      },
      "filter": {
        "url_ngram": {
          "type": "ngram",
          "min_gram": 3,
          "max_gram": 15
        }
      },
      "analyzer": {
        "url_truncated_ngram": {
          "type": "custom",
          "char_filter": ["truncate_url"],
          "tokenizer": "keyword",
          "filter": ["url_ngram"]
        }
      }
    }
  },
  "mappings": {
    "properties": {
      "url": {
        "type": "text",
        "analyzer": "url_truncated_ngram",
        "fields": {
          "raw": {
            "type": "keyword"
          }
        }
      }
    }
  }
}

Breakdown of the Configuration:

  • truncate_url char filter: Takes the first 100 characters of the URL and discards the rest. Perfect for your use case since meaningful parts like domain and path are almost always in the first 100 chars.
  • url_ngram token filter: Generates n-grams from 3 to 15 characters, but only from the truncated 100-char string.
  • url_truncated_ngram analyzer: Uses the keyword tokenizer (so the entire truncated URL is treated as a single token) before applying the n-gram filter—this ensures we get all possible 3-15 character substrings from the truncated URL.
  • url.raw field: Optional but useful if you ever need to retrieve or filter on the full original URL.

2. Test the Setup

Let's index your sample URL:

POST /url_index/_doc
{
  "url": "https://click.example.com/foo/bar/55gft/?qs=1952934d0ee8e2368ec7f7a921e3c6202b39365b9a2d26774c8122b8555ca21fce9d2344fc08a8ba40caede5e6901a112c6e89ead40892109eb8290d70571eab"
}

Now search for your target fragments:

Search for "example.com"

POST /url_index/_search
{
  "query": {
    "match": {
      "url": "example.com"
    }
  }
}

This will return the document, since "example.com" is within the first 100 characters of the URL.

Search for "/foo/bar"

POST /url_index/_search
{
  "query": {
    "match": {
      "url": "/foo/bar"
    }
  }
}

This also returns the document, as expected.

Bonus: Verify Truncation Works

If you search for a fragment from the long query string (like "1952934d"), it won't return the document—because that part is beyond the first 100 characters, so it was truncated before n-grams were generated. That's exactly the performance optimization you want!

Why This Works

By truncating first, we drastically reduce the number of n-grams generated. For your sample URL, the truncated version is:

https://click.example.com/foo/bar/55gft/?qs=1952934d0ee8e2368ec7f7a921e3c6202b39365b9a2d

Only this string is used to generate n-grams, not the full 200+ character URL. This cuts down on index size and speeds up both indexing and searching.

内容的提问来源于stack exchange,提问作者user2317003

相关产品推荐
方舟 Agent Plan

超全模态模型 × Harness 升级,最新支持 Deepseek-V4.1-Flash、GLM-5.3 系列、Doubao-Seedream-5.0-pro、Kimi-K3 (部分), 限时 9.9 元起

最近更新时间:2026.04.30 13:02:41